TRACE: reward hacking in RL fine-tuning
A controlled study of reward hacking: SmolLM2-360M trained with SFT then GRPO against a weak and a strong verifier, with logit lens and probes to find where the two models diverge.
A controlled study of reward hacking: SmolLM2-360M trained with SFT then GRPO against a weak and a strong verifier, with logit lens and probes to find where the two models diverge.
A five-agent word-puzzle pipeline (writer, judge, improver, finalizer, explanation) with bounded retries and cached variants per level.
LLM-as-judge harness across four model pairs, calibrated against blind human ratings.
A live no-signup app for planning where a group eats: 97K+ place profiles across ~200 UK towns, semantic search on BGE-small, under 400ms per query on free-tier infrastructure.
A fixed-size, model-agnostic embedding representation (0.955 MRR) and a ~1K-parameter memory scorer that matches a cross-encoder reranker with ~278,000× fewer parameters. Pre-prints in progress.
A TikTok-style recommender built from scratch: scoring, Gemini embeddings, FAISS indexing and the exploitation-exploration trade-off.
Undergraduate dissertation: federated training across 3 clients on Ethereum transactions with 90%+ class imbalance, reaching 99.7% AUC-ROC.
An interactive tool for seeing regression: tweak slope, intercept, polynomial order and learning rate, and watch L1 and L2 fits update in real time.
Predicting and segmenting support for climate action across 73,000+ respondents in 73 countries: 89.76% AUC-ROC with a random forest, six country typologies with k-means.