Can a model learn from data it never sees?

This is my undergraduate dissertation project: a federated learning system for detecting fraudulent Ethereum transactions. Three simulated clients train a shared model for 10 rounds, and none of them ever sends its raw data anywhere. This page walks through the design, the results, and the honest problems we found once we looked closely at those results.

The setup, in one sentence

3 clients, each holding a private slice of Ethereum transaction data, train one shared fraud-detection model for 10 rounds without ever exchanging that data.

Each client trains locally, then sends only its model's weights to a central server. The server averages the weights across clients and sends the improved model back. This is federated learning: the model travels, the data never does.

Federated modelfinal round, held-out test set
AUC-ROC: 0.9970 F1: 0.9814 Precision: 0.9976 Recall: 0.9656 Accuracy: 0.9901

These numbers are near perfect for a fraud detector, and that is not a good sign by itself. The rest of this page explains what we think produced them, and why we do not fully trust them.

The problem

Decentralised finance (DeFi) lets people transact without a bank or broker in the middle. That is also what makes fraud detection hard: there is no single institution holding everyone's transaction history to check for suspicious patterns.

The obvious fix, pooling everyone's data into one place to train a fraud model, breaks the point of DeFi and creates a single target for a breach. We wanted to know whether a set of independent platforms could train one shared fraud detector together, each keeping its own transactions private the whole time.

Privacy

What it shows
Centralising transaction data for risk scoring means one breach exposes everyone.

Fit with DeFi

What it shows
DeFi is built to be distributed and trustless. A central data store works against that.

Single point of failure

What it shows
Whoever holds the pooled data becomes the one entity that, if compromised, brings the whole system down.

The data

We used a public dataset of anonymised Ethereum transactions from Kaggle, holding transaction values, gas prices, timing patterns, and smart contract interaction features. Fraudulent transactions were rare: under 1% of the dataset, which is typical for fraud detection and also the reason a model cannot just be graded on accuracy.

A model trained on that split can reach 99% accuracy by labelling every transaction as legitimate, and be useless. So we watched precision, recall, F1 and AUC-ROC instead, and we used SMOTE (Synthetic Minority Over-sampling Technique), a method that manufactures new fraud examples by interpolating between real ones, to give the model enough fraud cases to learn from during training.

legitimate, over 99% Fraudulent transactions: under 1% fraud, under 1% ↑ Share of transactions
Fraud is a sliver of the dataset. The dissertation reports fraudulent transactions at under 1% of the total, which is why we used SMOTE to synthesise extra fraud examples for training rather than train on the raw split.
SMOTE has a cost we did not fully account for at the time. A synthetic fraud example is an interpolation between two real ones, so it can be easier to classify than a genuine, novel fraud pattern would be. See how honest are these numbers.

How it works

One central server coordinates three clients. Each client is a stand-in for a DeFi platform: it holds its own private data and its own copy of the model. The server never sees a client's data, only the numbers that describe how that client's copy of the model changed.

FL server model aggregation global model updates back Client 1 local data local model Client 2 local data local model Client 3 local data local model raw data never leaves a client
The server only ever exchanges model parameters. It sends the current global model's weights down to each client (solid arrows), and each client sends its updated weights back (dashed arrow, shown for one client to keep the diagram readable, but every client does the same). No transaction data crosses this boundary in either direction.

Each communication round follows the same five steps. We ran 10 rounds in total, using the Flower framework to orchestrate the clients and server, and FedAvg (federated averaging) to combine their updates: each client's contribution to the new global model is weighted by how much data that client trained on.

Distribute

The server sends the current global model's weights to all 3 clients.

Train locally

Each client trains the model on its own private data for 5 epochs, batch size 32.

Send updates

Clients send back only their updated weights, never raw transactions.

Aggregate

The server averages the updates with FedAvg, weighted by each client's dataset size.

Repeat

The improved global model goes back out. This loop ran for 10 rounds.
Raw transaction data never leaves a client's own machine at any step. Only model weights travel, in both directions.

The model

The model itself is a small multi-layer perceptron (MLP), a neural network of a few fully connected layers, called FraudDetectionNet: two hidden layers of 64 and 32 neurons with ReLU activations, a dropout layer for regularisation, and a single sigmoid output that scores how likely a transaction is to be fraudulent. All numeric features were standardised with StandardScaler, fitted once and reused across all clients so every client scaled its features the same way.

FraudDetectionNet(
  input -> Linear(64) -> ReLU -> Dropout(0.2)
        -> Linear(32) -> ReLU -> Dropout(0.2)
        -> Linear(1)  -> Sigmoid
)
# small on purpose: easier to debug and iterate across 3 clients and 10 rounds

Code layout

The project splits into six small files. Each client runs the same clients.py code on its own data; only server.py and evaluate.py ever see anything aggregated.

data_split.py

Splits the Ethereum dataset across the 3 simulated clients and saves the fitted feature columns, so every client encodes its inputs the same way.

model.py

Defines FraudDetectionNet, the shared architecture every client and the server load.

clients.py

The FL client class: trains locally each round, evaluates on its own test split, and saves that client's metrics to a JSON file.

server.py

Runs the FL server: distributes the global model, waits for client updates, aggregates them with FedAvg.

evaluate.py

Trains and scores the centralised baseline on the pooled dataset, for comparison with the federated result.

app.py

A small Flask dashboard, reading the saved metrics JSON files to show training progress round by round.

Results

By the final round, the global model scored close to the ceiling on every metric we tracked. We also trained a centralised model on the same data pooled together, as a baseline. The confusion matrix below, counted on the same global test set, gives an exact comparison rather than a description: of 436 real fraud cases, the federated model missed 15 and the centralised model missed 23. Both raised exactly one false alarm on a legitimate transaction.

Confusion matrix

A confusion matrix counts every prediction against the true label. The diagonal (shaded mint) is where the model got it right; the off-diagonal cells (shaded rose) are the two ways it can get it wrong: a false alarm, or a missed fraud.

Federated modelglobal test set, 1,969 transactions
Predicted normalPredicted fraud
True normal1,5321
True fraud15421
Missed 15 of 436 fraud cases, and raised 1 false alarm on the 1,533 legitimate ones.
Centralised modelsame test set, 1,969 transactions
Predicted normalPredicted fraud
True normal1,5321
True fraud23413
Missed 23 of 436 fraud cases, and also raised exactly 1 false alarm.
MetricFederatedCentralisedComputed as
Recall (fraud)96.56%94.73%caught fraud ÷ all real fraud
Precision (fraud)99.76%99.76%caught fraud ÷ all fraud alarms
F198.14%97.18%balance of the two above
Accuracy99.19%98.78%all correct ÷ all transactions
Recall, precision and F1 computed from these counts match the headline numbers elsewhere on this page almost exactly. Accuracy computed here (99.19%) is close to, but not identical to, the 99.01% quoted from the dissertation text, most likely because that figure came from a slightly different evaluation run than the one this confusion matrix was drawn from.
100% 75% 50% 25% 0% AUC-ROC: 99.70% 99.7% AUC-ROC F1: 98.14% 98.14% F1 Precision: 99.76% 99.76% Precision Recall: 96.56% 96.56% Recall Accuracy: 99.01% 99.01% Accuracy
Every headline metric lands above 96%, with precision at 99.76% and recall at 96.56%. High precision means the model rarely raises a false alarm on a legitimate transaction; high recall means it catches most of the real fraud. Both together, this cleanly, is unusual for fraud detection, and it is the reason we do not take these numbers at face value. See how honest are these numbers.

Progress across the 10 rounds

The chart below tracks four metrics on the global model as training moved through its 10 rounds. These values are read off the original training plots rather than an exact log, so treat them as approximate: the shape of the curve is the point, not the third decimal place.

100% 75% 50% 25% 0% AUC-ROC AUC-ROC Accuracy Accuracy F1 F1 Recall Recall 0 2 4 6 8 10 Federated learning round
All four metrics jump from near chance at round 0 to above 90% by round 2, then stay roughly flat for the remaining eight rounds. Loss follows the mirror shape: it fell from about 0.70 at round 0 to about 0.07 by round 2, then hovered between 0.07 and 0.11 for the rest of training. A model that reaches its ceiling this fast, this early, and stays completely flat afterwards is one of the signs that made us suspicious of the final numbers.

Client by client

The global metrics above are an average across clients. Looking at each client on its own tells a more interesting story: two of the three clients scored essentially perfectly, and the third did not.

ClientFinal training lossAUC-ROCNote
Client 10.00157100%Perfect score across every metric.
Client 20.00140100%Also perfect across every metric.
Client 3higher97.87%Noticeably below the other two, and the only client that behaves like we would expect.
Two clients scoring a perfect 100% on every metric, with training loss near zero, is not a sign of a strong model. In real machine learning this almost always means the evaluation setup let information leak from the test data into training, or that the test set was too small or too easy to be a fair check.

How honest are these numbers

The dissertation's own reflection on this project is more useful than the metrics themselves, so we keep it here in full rather than smoothing it over. Three things sit behind the near-perfect scores above.

What looks good

  • The federated model trained successfully without any client's data ever leaving that client.
  • Performance stayed competitive with a centralised baseline, which is the core claim the project set out to test.
  • The model improved consistently across the first few rounds, which is what a working aggregation process should look like.

What is concerning

  • Client 1 and client 2 both hit 100% on every metric, with training loss near zero. Real fraud detectors do not do this.
  • The final loss curve went flat almost immediately and stayed flat, a pattern more consistent with overfitting than genuine learning.
  • 99%+ precision alongside 96%+ recall at the same time is rare in production fraud systems.

Three likely causes

SMOTE made the task easier

What it does
Generates synthetic fraud examples by interpolating between real ones, to fix the under-1% class imbalance.
What it might have caused
The model may have learned to recognise the interpolation pattern itself, rather than real fraud signatures. A synthetic example is mathematically closer to its training neighbours than a genuinely new fraud case would be.

Possible data leakage

What it does
If SMOTE or the StandardScaler were fit before the train/test split, information from the test set leaks backward into training.
What it might have caused
Test metrics that look far better than the model would actually achieve on truly unseen transactions.

One dataset split three ways

What it does
All three clients drew from the same underlying Ethereum dataset, just partitioned, rather than from genuinely different institutions.
What it might have caused
Data that is more similar across clients (closer to IID, independent and identically distributed) than real DeFi platforms would ever be, which flatters federated averaging.
LessonA model that appears to hit its ceiling instantly and never moves again is a signal to check the evaluation, not a result to publish. We are reporting these numbers here as data points from a proof-of-concept, not as evidence that this system is production-ready.

Limits

This was a proof of concept, not a deployment. The goal was to show federated learning could plausibly work for DeFi fraud detection, not to build a system ready for real transactions.

Three clients from one dataset is not three institutions. Real DeFi platforms would differ far more in their user populations, transaction patterns and fraud types than a single dataset split three ways can capture.

The evaluation used one time-blind split. Fraud patterns shift over time. Training and testing on the same period, rather than training on older data and testing on newer, cannot show whether the model would generalise to fraud it has never seen the shape of.

Near-perfect scores are themselves a limit. We cannot rule out SMOTE-driven leakage or an easy test split from the information available in the original write-up, so the exact size of the true performance gap is unknown.

What we would do differently

Looking back, the next steps would be:

  1. Split training and test data by time period, so the model is evaluated on fraud patterns it has genuinely not seen before.
  2. Replace SMOTE with class-weighted loss, which penalises missed fraud more heavily without inventing synthetic transactions.
  3. Hold out a validation set that is never touched until the very end of the project.
  4. Fit the scaler and any resampling only on the training fold, after the split, to remove the leakage risk entirely.
  5. Simulate clients from genuinely different data sources, or at least different time windows, instead of one dataset partitioned three ways.
  6. Add differential privacy to each client's updates, so the privacy guarantee is mathematical, not just "we did not send raw data".

Stack: Python, PyTorch, the Flower federated learning framework, scikit-learn, pandas, NumPy, and a small Flask app used to view training metrics during development.