Can five small LLM calls write a puzzle that actually teaches a word?

Vocabscapes is a word puzzle game. Every level's clue sentence is written on demand by a chain of five LangGraph agents: a writer, a judge, an improver, a finalizer and an explainer. A bounded retry loop keeps the judge from rejecting forever, and a small Postgres cache means most levels are served, not regenerated.

Live at vocabscapes.web.app

A seed word entering the pipeline

The engine reads a seed word for the level (one line from a 6,559-word list) and finds every shorter word that can be spelled using only its letters, no letter used more times than it appears in the seed. This is a worked example built from the project's own code and prompts, not a captured production run.

seed word: GARDEN valid subwords (partial): DANGER, RANGE, GRAND, ANGER, DEAR, NEAR, GEAR, DRAG, RAGE

This word list, plus a difficulty profile for the level number, is everything the writer agent sees.

writerfirst draft

"The GARDEN hid DANGER among the thorny roses."
target words: [DANGER]

judgeFAILED

Target word "DANGER" is not shown in action. A reader
can't tell what danger means just from this sentence.

improverrevision 1

"Every step past the fence risked real DANGER, since
the old bridge could collapse under a single foot."

finalizermasked for the player

Every step past the fence risked real [ _ _ _ ] (6),
since the old bridge could collapse under a single foot.
target word

The blank is always drawn as three underscores. The number in brackets is the real answer length, taken from len(t) in finalize_sentence, so a 6-letter answer still shows three marks.

The pipeline

Every level is generated by one LangGraph graph: a small state machine where each node is one call to Gemini. The graph in sentence_agent.py runs five agents in sequence, with one loop.

Writer

Picks a target word from the list, then writes a sentence that shows the word's meaning through action, not a dictionary-style definition.

Judge

Runs the sentence through 12 checks: does the target word actually appear, is it used correctly, is there enough context to guess it, is the grammar natural. Returns PASSED or FAILED with a reason.

Improver

Gets the failed sentence and the judge's reason, and writes a full replacement, not a patch.

the judge scores the improver's new sentence again, up to 5 times, before the graph moves on regardless of the verdict

Finalizer

Masks each target word in the sentence with a blank and stores the answer length. No model call, plain string code.

Explainer

Looks up each target word in a free dictionary API, asks Gemini which definition fits this sentence, and asks for a plain-English phonetic spelling and a simplified rewrite of the sentence.

All five calls use the same small model, gemini-flash-lite-latest, run at a high temperature (2.0, versus the library default of 1.0) so the writer does not produce the same sentence twice for the same word list. That variety is only safe because the judge and improver sit downstream of it and can reject a bad roll.

The difficulty curve

The game plans for 300 or more levels drawn from a roughly 6,000-word seed pool. Rather than one difficulty setting, difficulty_rating.py changes exactly one variable at a time as the level number climbs, so a player never feels two things get harder at once. Every value below is quoted directly from that file.

LevelsTarget wordsWord lengthSentence lengthVocabulary guidance
1 to 513 letters6 to 10 wordseveryday, a 10-year-old's words
10 to 2923 to 5 letters8 to 12 wordscommon household, nature, daily life
60 to 6914 to 6 letters10 to 14 wordsless common, still concrete and visual
250 to 25934 to 8 letters18 to 23 wordsphilosophical, technical, specialized
300 and up34 to 8 letters24 to 30 wordsunrestricted: archaic, technical, poetic

Word length is capped at 8 letters for the whole game, a fixed line in the file's own header comment. Each of these settings is text handed to the writer and improver prompts. Nothing in the code measures or rejects a sentence for missing the target length, so the numbers are guidance the model can drift from, not a hard constraint the judge checks.

Caching and retries

Calling an LLM five times for every level a player opens does not scale. So the backend treats a level as content, not a live computation: once a word list produces a sentence that the judge marks PASSED, it is written to a level_data table in Postgres, keyed by level number.

The next time anyone requests that level number, generate_by_level first asks the database for existing passed rows. If two or more already exist, it picks one at random and serves it, no Gemini call at all. Only when fewer than two passed variants exist does it draw a fresh seed word, run the anagram engine, and invoke the full five-agent graph.

Serve from cache the common path

if len(level_ids) >= 2:
    random_id = random.choice(level_ids)
    return LevelData.fetch_level(random_id)

Two passed variants per level means repeat visits, and repeat players on the same level number, see some variety without a new generation call each time.

Bounded retries inside the graph

if verdict == "FAILED" and cycle_count <= 5:
    return "improve"
return "finalize"

A sentence the judge keeps rejecting still reaches the player after 5 improve cycles, marked with whatever verdict it last got. A "report level" endpoint lets a player flag a bad one, which flips it to FAILED so it drops out of the cache.

A level is not deleted when it fails the judge or gets reported: the row stays, only the critc_verdict flag changes, and the cache query filters on critc_verdict = PASSED.

Engineering decisions

Two choices shape most of the pipeline's behaviour: why generation is a chain of agents instead of one prompt, and why the game caches text instead of building it live.

One critique call, not a bigger writer prompt

The writer prompt is already long: show-don't-tell instructions, banned words, proper-noun rules, worked examples. Folding the judge's 12 checks into the same call would make the model grade its own homework in one pass. Splitting write and critique into separate calls means the second call reads the sentence fresh, the way a player would.

Store text, not call an API on every load

A level number is requested by every player who reaches it. Generating a fresh sentence each time would mean unpredictable latency and cost tied to Gemini, for content that does not need to change per player. Caching two passed variants trades a little repetition for flat, near-instant reads on the common path.

The word source

The anagram engine that produces each level's word list is a plain counting check, no LLM involved: a word is valid for a seed if every letter it uses appears at least as many times in the seed. It runs once against an 82,098-word UK English dictionary loaded into memory, and returns everything under about 8 letters that qualifies.

Below is a small, direct port of that function. It runs against a 40-word demo list embedded in this page, not the project's real 82,098-word file, so try a longer seed word to see more matches.

Try the subword check


matches in the demo list
0

What is not done yet

The next steps would be:

  1. Enforce the length and vocabulary settings after generation, not only inside the prompt, so drift gets caught before a level is cached.
  2. Feed reported levels back into the judge prompt as negative examples, instead of only flipping their verdict.