TRACE: reward hacking in RL fine-tuning
A controlled study of reward hacking: SmolLM2-360M trained with SFT then GRPO against a weak and a strong verifier, with logit lens and probes to find where the two models diverge.
A controlled study of reward hacking: SmolLM2-360M trained with SFT then GRPO against a weak and a strong verifier, with logit lens and probes to find where the two models diverge.
A five-agent word-puzzle pipeline (writer, judge, improver, finalizer, explanation) with bounded retries and cached variants per level.
LLM-as-judge harness across four model pairs, calibrated against blind human ratings.
A TikTok-style recommender built from scratch: scoring, Gemini embeddings, FAISS indexing and the exploitation-exploration trade-off.
Fundamental ML concepts with examples and explanations that feel like a conversation, not a lecture. A guide for building your machine learning foundation before you get to know about different...
Learning how VPC Peering lets different departments communicate directly without submitting tickets every time they need to access shared resources. Secure, private, and way more efficient.
Helping a bank fix their connectivity issues by setting up VPC Internet Gateways. Turns out, VPCs don't come with internet access by default - you gotta configure it yourself.
Using AWS Price Calculator to estimate costs for a surf shop's cloud migration. Played around with instance types, scaling, and discovered that cloud bills can get wild real quick