Can a model learn from data it never sees?
This is my undergraduate dissertation project: a federated learning system for detecting fraudulent Ethereum transactions. Three simulated clients train a shared model for 10 rounds, and none of them ever sends its raw data anywhere. This page walks through the design, the results, and the honest problems we found once we looked closely at those results.
- The problem
- The data
- How it works
- The model
- Results
- Client by client
- How honest are these numbers
- Limits
- What we would do differently
The setup, in one sentence
Each client trains locally, then sends only its model's weights to a central server. The server averages the weights across clients and sends the improved model back. This is federated learning: the model travels, the data never does.
These numbers are near perfect for a fraud detector, and that is not a good sign by itself. The rest of this page explains what we think produced them, and why we do not fully trust them.
The problem
Decentralised finance (DeFi) lets people transact without a bank or broker in the middle. That is also what makes fraud detection hard: there is no single institution holding everyone's transaction history to check for suspicious patterns.
The obvious fix, pooling everyone's data into one place to train a fraud model, breaks the point of DeFi and creates a single target for a breach. We wanted to know whether a set of independent platforms could train one shared fraud detector together, each keeping its own transactions private the whole time.
Privacy
- What it shows
- Centralising transaction data for risk scoring means one breach exposes everyone.
Fit with DeFi
- What it shows
- DeFi is built to be distributed and trustless. A central data store works against that.
Single point of failure
- What it shows
- Whoever holds the pooled data becomes the one entity that, if compromised, brings the whole system down.
The data
We used a public dataset of anonymised Ethereum transactions from Kaggle, holding transaction values, gas prices, timing patterns, and smart contract interaction features. Fraudulent transactions were rare: under 1% of the dataset, which is typical for fraud detection and also the reason a model cannot just be graded on accuracy.
A model trained on that split can reach 99% accuracy by labelling every transaction as legitimate, and be useless. So we watched precision, recall, F1 and AUC-ROC instead, and we used SMOTE (Synthetic Minority Over-sampling Technique), a method that manufactures new fraud examples by interpolating between real ones, to give the model enough fraud cases to learn from during training.
How it works
One central server coordinates three clients. Each client is a stand-in for a DeFi platform: it holds its own private data and its own copy of the model. The server never sees a client's data, only the numbers that describe how that client's copy of the model changed.
Each communication round follows the same five steps. We ran 10 rounds in total, using the Flower framework to orchestrate the clients and server, and FedAvg (federated averaging) to combine their updates: each client's contribution to the new global model is weighted by how much data that client trained on.
Distribute
The server sends the current global model's weights to all 3 clients.Train locally
Each client trains the model on its own private data for 5 epochs, batch size 32.Send updates
Clients send back only their updated weights, never raw transactions.Aggregate
The server averages the updates with FedAvg, weighted by each client's dataset size.Repeat
The improved global model goes back out. This loop ran for 10 rounds.The model
The model itself is a small multi-layer perceptron (MLP), a neural network of a few fully connected layers, called FraudDetectionNet: two hidden layers of 64 and 32 neurons with ReLU activations, a dropout layer for regularisation, and a single sigmoid output that scores how likely a transaction is to be fraudulent. All numeric features were standardised with StandardScaler, fitted once and reused across all clients so every client scaled its features the same way.
FraudDetectionNet(
input -> Linear(64) -> ReLU -> Dropout(0.2)
-> Linear(32) -> ReLU -> Dropout(0.2)
-> Linear(1) -> Sigmoid
)
# small on purpose: easier to debug and iterate across 3 clients and 10 rounds
Code layout
The project splits into six small files. Each client runs the same clients.py code on its own data; only server.py and evaluate.py ever see anything aggregated.
data_split.py
Splits the Ethereum dataset across the 3 simulated clients and saves the fitted feature columns, so every client encodes its inputs the same way.model.py
DefinesFraudDetectionNet, the shared architecture every client and the server load.clients.py
The FL client class: trains locally each round, evaluates on its own test split, and saves that client's metrics to a JSON file.server.py
Runs the FL server: distributes the global model, waits for client updates, aggregates them with FedAvg.evaluate.py
Trains and scores the centralised baseline on the pooled dataset, for comparison with the federated result.app.py
A small Flask dashboard, reading the saved metrics JSON files to show training progress round by round.Results
By the final round, the global model scored close to the ceiling on every metric we tracked. We also trained a centralised model on the same data pooled together, as a baseline. The confusion matrix below, counted on the same global test set, gives an exact comparison rather than a description: of 436 real fraud cases, the federated model missed 15 and the centralised model missed 23. Both raised exactly one false alarm on a legitimate transaction.
Confusion matrix
A confusion matrix counts every prediction against the true label. The diagonal (shaded mint) is where the model got it right; the off-diagonal cells (shaded rose) are the two ways it can get it wrong: a false alarm, or a missed fraud.
| Predicted normal | Predicted fraud | |
|---|---|---|
| True normal | 1,532 | 1 |
| True fraud | 15 | 421 |
| Predicted normal | Predicted fraud | |
|---|---|---|
| True normal | 1,532 | 1 |
| True fraud | 23 | 413 |
| Metric | Federated | Centralised | Computed as |
|---|---|---|---|
| Recall (fraud) | 96.56% | 94.73% | caught fraud ÷ all real fraud |
| Precision (fraud) | 99.76% | 99.76% | caught fraud ÷ all fraud alarms |
| F1 | 98.14% | 97.18% | balance of the two above |
| Accuracy | 99.19% | 98.78% | all correct ÷ all transactions |
Progress across the 10 rounds
The chart below tracks four metrics on the global model as training moved through its 10 rounds. These values are read off the original training plots rather than an exact log, so treat them as approximate: the shape of the curve is the point, not the third decimal place.
Client by client
The global metrics above are an average across clients. Looking at each client on its own tells a more interesting story: two of the three clients scored essentially perfectly, and the third did not.
| Client | Final training loss | AUC-ROC | Note |
|---|---|---|---|
| Client 1 | 0.00157 | 100% | Perfect score across every metric. |
| Client 2 | 0.00140 | 100% | Also perfect across every metric. |
| Client 3 | higher | 97.87% | Noticeably below the other two, and the only client that behaves like we would expect. |
How honest are these numbers
The dissertation's own reflection on this project is more useful than the metrics themselves, so we keep it here in full rather than smoothing it over. Three things sit behind the near-perfect scores above.
What looks good
- The federated model trained successfully without any client's data ever leaving that client.
- Performance stayed competitive with a centralised baseline, which is the core claim the project set out to test.
- The model improved consistently across the first few rounds, which is what a working aggregation process should look like.
What is concerning
- Client 1 and client 2 both hit 100% on every metric, with training loss near zero. Real fraud detectors do not do this.
- The final loss curve went flat almost immediately and stayed flat, a pattern more consistent with overfitting than genuine learning.
- 99%+ precision alongside 96%+ recall at the same time is rare in production fraud systems.
Three likely causes
SMOTE made the task easier
- What it does
- Generates synthetic fraud examples by interpolating between real ones, to fix the under-1% class imbalance.
- What it might have caused
- The model may have learned to recognise the interpolation pattern itself, rather than real fraud signatures. A synthetic example is mathematically closer to its training neighbours than a genuinely new fraud case would be.
Possible data leakage
- What it does
- If SMOTE or the StandardScaler were fit before the train/test split, information from the test set leaks backward into training.
- What it might have caused
- Test metrics that look far better than the model would actually achieve on truly unseen transactions.
One dataset split three ways
- What it does
- All three clients drew from the same underlying Ethereum dataset, just partitioned, rather than from genuinely different institutions.
- What it might have caused
- Data that is more similar across clients (closer to IID, independent and identically distributed) than real DeFi platforms would ever be, which flatters federated averaging.
Limits
This was a proof of concept, not a deployment. The goal was to show federated learning could plausibly work for DeFi fraud detection, not to build a system ready for real transactions.
Three clients from one dataset is not three institutions. Real DeFi platforms would differ far more in their user populations, transaction patterns and fraud types than a single dataset split three ways can capture.
The evaluation used one time-blind split. Fraud patterns shift over time. Training and testing on the same period, rather than training on older data and testing on newer, cannot show whether the model would generalise to fraud it has never seen the shape of.
Near-perfect scores are themselves a limit. We cannot rule out SMOTE-driven leakage or an easy test split from the information available in the original write-up, so the exact size of the true performance gap is unknown.
What we would do differently
Looking back, the next steps would be:
- Split training and test data by time period, so the model is evaluated on fraud patterns it has genuinely not seen before.
- Replace SMOTE with class-weighted loss, which penalises missed fraud more heavily without inventing synthetic transactions.
- Hold out a validation set that is never touched until the very end of the project.
- Fit the scaler and any resampling only on the training fold, after the split, to remove the leakage risk entirely.
- Simulate clients from genuinely different data sources, or at least different time windows, instead of one dataset partitioned three ways.
- Add differential privacy to each client's updates, so the privacy guarantee is mathematical, not just "we did not send raw data".
Stack: Python, PyTorch, the Flower federated learning framework, scikit-learn, pandas, NumPy, and a small Flask app used to view training metrics during development.