Can you trust the judge grading your puzzle generator?
Vocabscapes generates word-puzzle sentences with one model and grades them with another, in a write→critique loop. This project measures whether that grader can be trusted, by checking it against blind human ratings — and finds that in one cell of the design, it was wrong by 39 points.
The setup
Vocabscapes is a word-puzzle game. Each level gives the player a word list and asks them to write one sentence that uses every word. A writer model drafts that sentence; a critic model grades it against a 12-point rubric and returns PASSED or FAILED per check. If it fails, the writer retries, up to 2 retries (3 attempts total) before the level is marked a Hard FAIL.
This project turns that write→critique loop into a benchmark. It compares two candidate models — Qwen3.8-27B and google/gemma-4-E2B-it — as both writer and critic, across a full 2×2 design: 140 levels × 4 (writer, critic) pairs, spanning 28 difficulty tiers. Two of the four pairs are same-model cells, where a model grades its own writing; the other two are cross-model.
The question this project actually answers is narrower and more useful than "which model writes better sentences." It is: can the critic's verdict be trusted at all, and if not, by how much and in which direction does it lie.
| Writer \ Critic | Qwen critic | Gemma critic |
|---|---|---|
| Qwen writer | qwen→qwen same-model | qwen→gemma cross-model |
| Gemma writer | gemma→qwen cross-model | gemma→gemma same-model |
Run 1 config Qwen3.8-27B-FP8 + gemma-4-E2B-it reasoning: on, both roles, both models token budget: 16384 (shared, uncapped split) sampling: per-model recipe from each model's card max attempts: 3 (2 retries) structured output: xgrammar, whitespace disabled
vllm serve processes.Pipeline
Word lists for the 140 sampled levels are generated live from Vocabscapes's own deterministic level generator — a seed word expanded into valid subwords — rather than read from a pre-existing CSV. The logged history only covered 18 of the 31 difficulty tiers, not enough to fill the 28 the sample needed, so the harness calls the generator directly for any level number.
Generate
Pull a level's word list from Vocabscapes's own level generator. 5 levels per tier × 28 tiers = 140 levels, spanning the full difficulty range.Write
The writer model drafts one sentence using every target word, for all 4 (writer, critic) cells.Critique
The critic model checks the sentence against the 12-point rubric and returns a per-check PASSED/FAILED verdict. Up to 2 retries on failure.Calibrate
Draw a stratified 80-row sample (20 PASSED + 20 FAILED per cell). A human and an isolated Claude subagent each re-grade it blind — no file access, no view of the critic's verdict or the cell label.Correct
Blend the critic's raw pass-rate with the blind raters' sensitivity and specificity into a bias-corrected rate, with a bootstrap confidence interval.The run logs every attempt to one of 4 separate results_<cell>.jsonl files, and every error — timeout, malformed structured output, dropped connection — to its own error log, never silently. A level whose call fails before a verdict is returned is abandoned, not a Hard FAIL: one is an infrastructure fault, the other is a result, and folding them together once made a broken run look like a strict critic (see §5).
What is judged
The critic checks 12 things about a candidate sentence: that it uses the seed word if required, that every target word actually appears, that the answer key is accurate, proper-noun handling, lexical integrity, substring overlap, acronym handling, output formatting, contextual richness, word efficiency, grammar, and semantic validity. A sentence passes only if checks 2–12 all pass — check 1 (seed word) is out of scope here because of a harness quirk described below.
Which checks are decidable, and which are judgment calls
- Checks 2 and 3 — target word existence, answer key accuracy — are exact word-boundary matching. They have one correct answer, checkable by a script.
- Checks 4, 6, 7, 8 — proper nouns, lexical integrity, acronyms, output style — are mechanical. Two independent raters agree on them almost perfectly (77–79 of 80 rows).
- Checks 5, 9, 10, 11, 12 — substring overlap, contextual richness, word efficiency, grammar, semantic validity — are judgment calls. Two raters agree on as few as 58 of 80 rows here.
Check 9, "could a player deduce the word from context," carries the single largest disagreement between raters: 22 of 80 rows. See §7.
Raw results: what the critic said
The critic's raw, uncorrected numbers. Everything downstream uses the attempt pass-rate p, not the level pass-rate — the pipeline retries until a level passes, so a level rate mostly answers "did retrying eventually work" and hides the critic's actual per-call behaviour. Level rates for this run sit at 84–98% and are close to useless for judging the critic.
| cell | attempts | attempt pass-rate p | level pass-rate | hard fail |
|---|---|---|---|---|
| qwen→qwen | 208 | 63.0% | 94.9% | 7 |
| qwen→gemma | 166 | 82.5% | 97.9% | 3 |
| gemma→qwen | 402 | 3.5% | 10.0% | 126 |
| gemma→gemma | 233 | 50.6% | 84.3% | 22 |
Calibration against blind human ratings
The critic is an unvalidated instrument. Its raw pass-rate p measures the writer's quality and the critic's own errors together, inseparably. Calibration pulls them apart.
A stratified 20-row sample per cell (80 rows total, half critic-PASSED and half critic-FAILED) is drawn from the real logged results — not fabricated separately. A human rater and an isolated Claude subagent each re-grade every row against the same 12-point rubric, blind: they see only the sentence, the word list, and the difficulty settings, never the critic's verdict or which cell produced the row.
The correction, in one paragraph
Of the rows the critic passed, the fraction the blind rater also passes gives a — the critic's precision on its own accepts. Of the rows the critic failed, the (rejection-weighted) fraction the blind rater would have passed gives b — how often the critic threw away a good sentence. Blending them by the critic's real attempt rate gives the corrected rate: theta = p·a + (1-p)·b. A 95% interval comes from resampling each pile 4,000 times.
| cell | a (passed pile) | b (failed pile) |
|---|---|---|
| qwen→qwen | 0.80 | 0.87 |
| qwen→gemma | 0.70 | 0.69 |
| gemma→qwen | 1.00 | 0.40 |
| gemma→gemma | 0.60 | 0.22 |
a and b are counts out of 10 rows per pile per cell, with b weighted by each level's own rejection count so a thrice-rejected level isn't undercounted.| cell | q1 sensitivity | q0 specificity |
|---|---|---|
| qwen→qwen | 0.61 | 0.28† |
| qwen→gemma | 0.83 | 0.18† |
| gemma→qwen | 0.08 | 1.00‡ |
| gemma→gemma | 0.74 | 0.66 |
q1 is a critic's hit-rate on genuinely good sentences, q0 on genuinely bad ones. † unstable — the interval spans nearly the whole 0–1 range. ‡ a ceiling forced by the sample (10/10), not a measured value. The one solid reading: the Gemma critic on Qwen writing never accepts a bad sentence (q0=1.00) but rejects roughly 12 good ones for every one it lets through (q1=0.08) — a usable description of a broken judge that holds under every rater.Corrected results
This is the number to quote: the human rater's verdicts, with the two objectively decidable checks (2 and 3) settled by a script instead of by opinion. Every pile has at least two successes, so the bootstrap behaves throughout.
| cell | critic p | corrected theta | 95% CI | shift, points |
|---|---|---|---|---|
| qwen→qwen | 0.630 | 0.825 | 0.63–0.98 | +19.5 |
| qwen→gemma | 0.825 | 0.698 | 0.45–0.91 | −12.7 |
| gemma→qwen | 0.035 | 0.421 | 0.13–0.71 | +38.6 |
| gemma→gemma | 0.506 | 0.411 | 0.20–0.62 | −9.5 |
The ranking survives; the scores don't
| rank | human + script | human raw | Claude |
|---|---|---|---|
| 1 | qwen→qwen 0.825 | qwen→qwen 0.825 | qwen→qwen 0.775 |
| 2 | qwen→gemma 0.698 | qwen→gemma 0.698 | qwen→gemma 0.615 |
| 3 | gemma→qwen 0.421 | gemma→gemma 0.598 | gemma→gemma 0.051 |
| 4 | gemma→gemma 0.411 | gemma→qwen 0.421 | gemma→qwen 0.014 |
Ranks 1 and 2 hold under all three raters and rubric variants. Ranks 3 and 4 sit within 0.01 in the headline and swap order between variants — treat them as tied. What survives every rater: Qwen writes better than gemma-4-E2B, by 0.28 to 0.68 depending on rater. That conclusion needs no rater to be correct, only self-consistent.
Where the rubric breaks
A second blind rater (a fresh Claude subagent, no file or internet access) re-graded the same 80 rows. Its purpose is a robustness check, not a second ground truth: it tells apart "the critic is genuinely miscalibrated" from "the rubric is too ambiguous for any rater to apply consistently." The answer turned out to be the second one.
Two competent raters, 80 identical rows, 30 rows apart
Human passes 55 of 80. Claude passes 31 of 80. The two agree on only 50 of 80 verdicts — Cohen's kappa 0.31, "fair" agreement, well short of what a published point estimate needs. The two blind raters agree with each other no better than either agrees with the instrument they were brought in to audit.
| check | agreement | human fails | Claude fails | kind |
|---|---|---|---|---|
| 8 output style | 79/80 | 0 | 1 | mechanical |
| 7 acronyms | 78/80 | 2 | 2 | mechanical |
| 6 lexical integrity | 77/80 | 6 | 7 | mechanical |
| 4 proper nouns | 77/80 | 0 | 3 | mechanical |
| 2 target word existence | 73/80 | 1 | 8 | decidable |
| 3 answer key accuracy | 68/80 | 2 | 14 | decidable |
| 5 substring overlap | 65/80 | 7 | 14 | mechanical |
| 12 semantic validity | 63/80 | 18 | 19 | judgement |
| 10 word efficiency | 63/80 | 15 | 18 | judgement |
| 11 grammar | 61/80 | 14 | 23 | judgement |
| 9 contextual richness | 58/80 | 17 | 31 | judgement |
Checks 2 and 3 stood out: both are exact word-boundary matching against the target words and word list, with one objectively correct answer — they should not have disagreed at all. Running a script over all 80 rows settled it: the human rater, not Claude, was careless there. Human wrong on 7 of 7 disputed rows for check 2, and 12 of 12 for check 3. Replacing those two checks with the script moved gemma→gemma from 0.598 to 0.411 and raised overall kappa from 0.31 to 0.37 — better, still only "fair." Adjudication fixes the mechanical disagreement and leaves the judgment disagreement, mostly check 9, completely intact.
Cost of finding this out
1,009 writer/critic attempt pairs, 2,080 model calls, 6h 27m wall time on one RTX PRO 6000 Blackwell GPU running both models concurrently. 7.89M tokens generated, 96.5% of it reasoning. Qwen is roughly 3× slower per token and emits roughly 6× more tokens per call than Gemma, so it accounts for 95% of the run's compute despite being one side of a 2×2 grid.
Only 1.25% of calls (26 of 2,080) hit the token budget before finishing — comfortably under the 10.8% truncation rate that forced an earlier budget increase from 8,192 to 16,384 tokens.
Conclusion
- The raw pass-rate cannot be reported as writer quality. Calibration moves every cell, and not in one direction — one pair rises 19.5 points, another falls 12.7, a third rises 38.6.
- Quote a corrected rate with its interval, never bare. The two Qwen-written cells are tight enough to be useful. The two Gemma-written cells land near 0.4 with intervals half the width of the scale, and swap order between rater variants — report them as "roughly 0.4, poorly resolved," not as a rank.
- The writer ranking is the one thing that survives every rater. Qwen beats gemma-4-E2B by 0.28 to 0.68 depending on which rater is used — without needing any single rater to be correct.
- The apparent same-model self-preference bias does not survive calibration. The raw +74.3-point gap for Gemma judging its own writing is mostly critic leniency; the two cells carrying that claim are exactly the two whose corrected rate is unstable across raters.
- The blocking problem is the rubric, not the sample size. Two competent raters applying the same 12-point rubric to the same 80 rows land 30 rows apart, almost entirely on the four checks whose wording is least operational (grammar, semantic validity, word efficiency, and above all "could a player deduce the word from context").