Here is an exercise from the October refill, exactly as a learner gets it. It practises reflexive pronouns at B1. Try it before reading on.

The interactive exercise needs JavaScript.

If you got gap 1 or gap 5 wrong, you probably chose mich for nehme … vor or mir for freue … auf. That is the point of the exercise. Dative versus accusative reflexive pronouns is exactly where B1 learners slip, so those are exactly the options that should tempt them. The correct answers practically write themselves. The wrong answers are where the teaching happens.

What a good wrong answer looks like

A distractor has three jobs at once. It has to be real German, it has to be wrong in this exact sentence, and it has to be wrong for the reason the exercise is about. Miss any one of them and the gap stops teaching. In practice there are four ways to miss.

One gap, four options. 'im' is correct: location takes the dative. 'in den' is a good distractor with one error inside the target grammar. 'in dem' is flawed because it is also standard German, a second right answer. 'bei den' is flawed because it is wrong in two independent ways.
One gap, four options. Only one of the three wrong answers is a good distractor. in dem Park is stiff but correct, so a careful learner who picks it is marked wrong for being right.

The double error is the subtle one. If a learner picks bei den, did they get the preposition wrong or the case? The exercise can’t tell them, and the explanation has to cover two rules at once. A good distractor isolates one decision.

The failure that dominated

When I re-checked every live exercise against the current standard, 696 of 829 failed somewhere. Sorting the failures by type made the priority obvious.

More than half of all failures had a second right answer. That number changed how I thought about generation. I had been treating distractors as the last, easy step after a nice text. They are the hardest part of the specification, and the place where a model’s fluency works against you. A fluent model knows that in dem Park is fine German, so it doesn’t see why it’s a bad option.

From rules for the checker to rules for the writer

The first versions of the new pipeline (V2, V3) caught bad distractors and repaired them. That worked, but it was expensive: most drafts needed at least one repair round. So I moved the same rules upstream into the generation prompt (V4). Every wrong option must differ from the answer only in the grammar point being practised, must be wrong in one way only, and must be something a learner at that level would plausibly choose. The exercise as a whole must not be solvable by a surface pattern.

On 40 held-out tasks the share of first drafts passing every gap check rose from 8% to 32%. On a pre-registered comparison over 100 tasks per version, the share of accepted exercises that a blind audit graded flawed fell from 23% to 12%.

Pipeline versions: what V1–V6 mean
  1. V1
    Original pipelineuntil Sep 2026

    One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.

  2. V2
    Generate, verify, repair28 Sep 2026

    A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.

  3. V3
    + coherence check30 Sep 2026

    After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.

  4. V4
    + design rules in the prompt2 Oct 2026

    The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.

  5. V5
    + stronger writer, editor gate3 Oct 2026

    A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.

  6. V6
    + mixed-grammar tolerance4 Oct 2026

    Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.

Each version keeps everything before it. Charts and text name the version that produced a number.

The surface-pattern trapOne accepted exercise practised modal adverbs. Every distractor was an inflected form of an adverb that never inflects, such as hoffentliche. A learner could solve every gap by picking the only uninflected option, without knowing any grammar. Each gap passed every check. Only reading the whole exercise showed the problem, and it became an explicit rule for the writer and for the editor gate.

Two product rules that came from reading failures

Strict checks find real problems, but they also need product judgment about what actually helps a learner. Two of my rules came from reading rejected exercises, not from the checks.

Fixed phrases count as one choice. In an exercise on fixed verb phrases, in die Sprache and an die Sprache as distractors for zur Sprache change the preposition and the article together. A strict reading calls that a double error. But learners learn zur Sprache kommen as one unit, and zur is a fused form anyway. So the checks now treat the phrase as a single choice.

Up to 30% mixed-grammar gaps are tolerated (V6). Natural texts sometimes force a gap to touch a neighbouring grammar point. One such gap in seven doesn’t mislead anyone, and rejecting the whole exercise for it wastes good work. An exercise-wide surface pattern, like the adverb example above, stays fatal. On 40 confirmation tasks the rule changed accepted exercises from 30 to 31, flawed ones from 1 to 0, and the cost per clean exercise from $0.099 to $0.089.

Three options instead of four: tested, not adopted

telc uses three options per gap at A2 to B2, and fewer distractors means fewer chances for a second right answer. I tested it with the rule written in advance: switch only if the cost per clean exercise drops by at least 15%. It didn’t. It came out at $0.142 against $0.137 for four options. The reason was interesting: with the new prompt, first drafts already had a second right answer in only 2.2% of gaps, so there was little left to fix. Four options stay: practice that is a little harder than the actual exam does no harm.

If you generate assessment content

  1. Specify the distractors, not just the answer. Each should be wrong in one way, for the reason being taught.
  2. Count failure types before fixing anything. One type will dominate.
  3. Move rules from the checker into the writer once they work. Repairs are the expensive way to comply.
  4. Read whole exercises. Some defects only exist at the level of the exercise, never in a single gap.
  5. Let product judgment set the strictness. Some “errors” are what learners actually learn.