InfiniteGrammar.de sells one thing: a lot of practice on one narrow piece of German grammar. Every exercise is a short text with five to seven gaps, four options per gap and an explanation for each answer. Large language models write all of them. I’m the only person who could check them, and I can’t read thousands.

For most of 2026 the pipeline (V1) graded its own work. A model wrote an exercise, the same model reviewed it, and a failed draft was rewritten until it passed or ran out of attempts. About 84% passed. A separate checker pass then re-read the survivors and flagged roughly half of them, about a third of those for doubled words and stray punctuation that a regex could fix. It felt like a quality process.

In late September I built a stricter check for the property that matters most in this format: exactly one option per gap should be correct. I ran it over the live library. 411 of 1,186 exercises failed it, meaning a learner could defend a second answer in at least one gap. The old process had approved every one of them.

This is the story of the ten days that followed, and of what changed in production afterwards.

Define “good” before choosing a model

The fix started with a definition, not with a better model. I wrote down what makes an exercise publishable as five checks it must pass and four scores it receives:

  • G1 · Correct German once the right answers are filled in.
  • G2 · Exactly one right answer in every gap.
  • G3 · On topic: each wrong option differs from the answer only in the grammar point being practised.
  • G4 · A correct explanation that names the deciding rule.
  • G5 · Format rules that code can check for free, such as gap count or a capital letter that gives the answer away.

Naturalness, level fit, how tempting the distractors are and how well the text hangs together became scores from 0 to 5. They rank exercises but never reject one. That split mattered more than I expected. Early versions of the judge failed perfectly good German for sounding slightly informal, and moving style out of the pass/fail checks removed most of those false alarms.

Test the judge before trusting it

An LLM judge is a model like any other, so it needs its own evaluation. The first version of my verifier graded five official telc exam texts, each with its official answer key, and failed three of them. On the hardest check it also changed its verdict on one gap in ten when I simply ran it twice.

I rebuilt it against four kinds of test data and replaced the single “is there one right answer?” call with two independent solvers that must agree, plus a stronger adjudicator that only speaks when they don’t. That story has its own article. In short, the check now finds 92–93% of real second answers, is right about 80% of the time when it raises one, and flips on 3% of gaps instead of 10%.

Change the shape of the pipeline

With a judge I could trust, I compared four ways of producing an exercise, using the same model on 30 held-out tasks. Rewriting a whole exercise after a failed review, which is what the old pipeline did, passed 7% of the time. Generating, verifying and then repairing only the gaps that failed passed 50%. That gap was larger than anything prompt tuning was likely to buy, so I postponed prompt optimisation and put the effort into the pipeline’s shape and the model.

The pipeline today: write the exercise, run free rule checks, two solvers and a judge decide whether each gap has exactly one right answer, failing gaps are repaired and re-checked, then a coherence check and an editor gate run before the exercise goes live. Anything still failing after its repair budget is stored as rejected and never served.
How an exercise reaches a learner today (V6). Repairs touch only the failing gaps. Every exercise that fails after its repair budget is kept for analysis but never shown to anyone.
G1
Correct German with the answers filled in
G2
Exactly one right answer per gap
G3
Wrong options differ only in the grammar being practised
G4
The explanation is right and names the deciding rule
G5
Format rules, such as a capital letter that gives the answer away
G6
The story holds together: who does what, who “she” is
FG1–FG6
The editor gate’s six issue types, from a second right answer to mixed grammar

Then I chose the model by what I actually pay for: cost per accepted exercise, not price per token. On 60 tasks the system had never seen, the old pipeline (V1) accepted 17% of attempts at $0.23 per accepted exercise. The new shape with a mid-sized model (V2) accepted 65% at $0.11.

Add what the checks couldn’t see

Reading accepted exercises by hand showed the next layer of problems. They weren’t in the gaps, they were in the story. In one, set in a shop, the shop assistant asked the customer, „Könnten Sie ihn bis heute Nachmittag reservieren?“ (“Could you reserve it until this afternoon?”). That’s the customer’s line, not the shop’s: the roles were reversed. In another, „diese Frau“ pointed at nobody. Every gap was fine and both exercises scored 0.95 out of 1.

So I added a coherence check as its own stage after the gap checks (V3), moved the rules for good wrong answers into the writing prompt (V4), and then added a stronger writer and a final read by the same model acting as a senior editor, with six named issue types (V5). Each addition had to win a comparison whose pass rule I wrote down before running it (how that worked). The final configuration (V6) accepted 31 of 40 fresh tasks. A blind audit graded 20 of those clean, 11 as having minor issues and none as flawed.

Pipeline versions: what V1–V6 mean
  1. V1
    Original pipelineuntil Sep 2026

    One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.

  2. V2
    Generate, verify, repair28 Sep 2026

    A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.

  3. V3
    + coherence check30 Sep 2026

    After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.

  4. V4
    + design rules in the prompt2 Oct 2026

    The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.

  5. V5
    + stronger writer, editor gate3 Oct 2026

    A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.

  6. V6
    + mixed-grammar tolerance4 Oct 2026

    Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.

Each version keeps everything before it. Charts and text name the version that produced a number.

Cost per accepted exercise
$0.23 →$0.058
all checks and repairs included
First drafts that pass the gap checks
8% →70%
before any repair
Cost per clean exercise
$0.089
accepted and graded clean

In production

On 5 October I re-checked all 829 live exercises against the new standard. 133 passed, and the rest went offline once replacements were ready (how I did that without breaking anyone’s progress). Over the following four days the replenishment pipeline made 1,379 attempts, accepted 974 of them and cost $78.14 in total. I read a random sample of 20 of the accepted ones: 14 clean, 6 with minor issues, none flawed.

Accepted in the refill
974
71% of 1,379 attempts
Cost per accepted exercise
$0.080
$78.14 for the whole refill
Random sample of 20
14 / 6 / 0
clean / minor / flawed

What I’d do differently

Build the test set before the first judge. I wrote a verifier, ran it on the live library and took 411 exercises offline on its word. Calibration later put that early check’s false-alarm rate at around 5–8%, so probably a few dozen good exercises went offline for nothing. Calibrate first, then act.

Write pass rules from day one. My first comparisons were judged after the fact. The later ones, with rules written before the run, were quicker to decide and much harder to argue with, including by me.

If you’re building with LLM output

  1. Write down “good” as checks before you choose a model, and keep matters of taste in scores rather than gates.
  2. Evaluate your judge on data with known answers, including examples that must pass.
  3. Measure cost per accepted, or better per clean, output. Price per token is the wrong unit.
  4. Repair the part that failed instead of regenerating everything.
  5. Keep reading the output yourself. Every check I added came from reading.
About the numbersHeld-out tests used 30–60 tasks each, and not always the same tasks, so compare directions more than decimals. Costs are at OpenAI Batch API prices. Audits were blind: the grader didn’t know which version produced an exercise. Sources are the project’s run reports and decision log from September and October 2026.