The most useful hour of the whole quality project was spent watching my verifier fail an exam.

The first version was running within a day. It had free format rules, a cheap model that solved each exercise like a learner, and a judge model that ruled on the German, the topic and the explanations. It looked sensible. Before letting it decide what learners see, I gave it five reading texts from official telc exam material, each with the official answer key. Professionals write and check those texts. A sound verifier should pass all five.

Mine passed two.

Three ways a judge goes wrong

The failures fell into three groups. I suspect most LLM-as-judge setups have all three.

It misses the cue that decides the answer. In a sentence built around „mit dessen Hilfe … war“, only one form fits because of a word several words earlier. The solver and the judge both read past it and agreed that a different form was acceptable. Two models agreeing is not the same as two models being right.

It confuses taste with error. The first judge failed exercises for German that a native editor would publish without a second look, because a phrase was slightly unusual. When a stronger model reviewed 30 of its “German error” flags, it called 27 of them style. I moved style into a 0–5 score that never blocks publication, and the share of German-error flags that were real errors rose from 43% to 83%.

It isn’t stable. Run the same check twice on the same exercise and a single solver changed its verdict on one gap in ten. A check that flips that often can’t be allowed to take content offline.

Four kinds of test data

I had no labelling team, so I built test data from what already existed. Each set answers a different question about the judge.

Four test sets: exam texts with official keys that should pass (5 telc texts, 50 gaps; every failure is a false alarm); mutations that should fail (288 copies with a planted defect; every pass is a miss); real defects from production (101 exercises, 62 confirmed defects); blind audits where graders don't know which version wrote the exercise (98 items became the gold set).
Four questions, four test sets. Exam texts measure false alarms, mutations measure misses, real defects measure what actually happens, and blind audits catch what no check measures yet.

The mutations were the cheapest to make and the most revealing. You take a good exercise and change one thing on purpose: replace a distractor with a second correct form, plant a grammar error such as mit die Freundin, or rewrite an explanation so it states a wrong rule. You know exactly what the right verdict is, so you can count misses without arguing about German. The exam texts do the opposite job. Nothing in them should fail, so every flag is a false alarm you can look at and learn from.

Two solvers and an adjudicator

The check that matters most is “exactly one right answer”. I tested several designs on 103 exercises with 641 audited gaps. The one that won is cheap at its core. Two independent runs of a small model (gpt-5-mini) solve every gap the way a learner would, with the options shuffled differently each time. If the two runs agree, that’s the verdict. If they disagree, a stronger model (gpt-5.2) adjudicates. Only about a quarter of exercises ever reach it.

On the exam texts, the new version passed three of five. The three gaps it still flags are ones I would happily argue about with a teacher, which is about as good as a check of this kind gets. I wrote them down as known limits rather than tuning the prompt until the test passed.

Verdicts that flip between identical runs
10% →3%
of gaps
Official exam texts passed
2 →3 of 5
every failure here is a false alarm
Cost of the full verifier
≈ $0.02
per exercise, at batch prices

Two other decisions came out of this. I did not build a panel of three judges, because the cheap solver pair already gave most of the benefit at a fraction of the price. And I kept matters of taste in scores. A judge that blocks content for style will block good content, and you won’t notice, because nobody complains about an exercise they never see.

A deliberately strict last gate

The last check, added in V5, is a frontier model (gpt-6.1-sol) reading the whole exercise like a senior editor and naming its problems with six issue types. I calibrated it on 98 exercises that had already been graded blind. It caught every exercise graded flawed, 17 of 17. It also rejected more than half of the ones graded clean.

I kept it strict on purpose. A wrongly rejected draft costs a few cents and gets regenerated. A wrong exercise teaches a learner something false and makes them doubt every other exercise they see. When the costs of the two errors are that lopsided, you tune for the expensive one, and you say so. Later I relaxed one specific rule the gate was too picky about, mixed-grammar gaps, rather than lowering the bar everywhere (V6, that decision is here).

Pipeline versions: what V1–V6 mean
  1. V1
    Original pipelineuntil Sep 2026

    One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.

  2. V2
    Generate, verify, repair28 Sep 2026

    A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.

  3. V3
    + coherence check30 Sep 2026

    After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.

  4. V4
    + design rules in the prompt2 Oct 2026

    The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.

  5. V5
    + stronger writer, editor gate3 Oct 2026

    A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.

  6. V6
    + mixed-grammar tolerance4 Oct 2026

    Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.

Each version keeps everything before it. Charts and text name the version that produced a number.

The mistake I made first

Before any of this calibration, I ran the first verifier over the live library and took 411 exercises offline because it flagged a second right answer. Later measurements put that early check’s false-alarm rate on real exercises at roughly 5–8%. So probably a few dozen good exercises disappeared for nothing. Nobody complained, which is exactly the problem with false rejections: they are invisible. The rule I work by now is simple. Calibrate a check, then let it act.

If you’re building an LLM judge

  1. Give it an exam with an answer key: real examples that must pass. False alarms are the error you won’t see otherwise.
  2. Plant known defects to count misses without debating each verdict.
  3. Separate taste from error. Gate on errors, score taste.
  4. Measure how often verdicts flip. If it’s high, pair or vote before you gate anything.
  5. Spend the expensive model only where the cheap ones disagree.
  6. Decide which error costs more, tune toward avoiding it, and write that down.
About the numbersLabels come from my own audits and from a stronger model’s review, not from a panel of native-speaker teachers. Precision and recall were measured on 103 exercises (641 gaps), the exam controls on 5 telc texts with 50 gaps, and the gate calibration on 98 blind-graded exercises.