I liked the distractor gate. It was a neat idea: after the gap checks pass, a strong model inspects every wrong option on its own and asks three questions. Is it wrong in one way only? Would a native speaker never say it here? Is it a real form a learner would choose? I wrote the prompt, ran it, and looked at the result.

It accepted 2 of 10 development tasks where the existing pipeline accepted 3, at twice the cost per accepted exercise. The pass rule I had written beforehand said the new version needed at least one more accepted exercise and no higher cost. It failed both conditions. The gate didn’t ship.

This article is about that habit: writing the rule before the test. It is the single practice that made evidence-based product decisions possible.

The setup

Three things made the comparisons trustworthy enough to act on.

Separate task sets. A task is a grammar section, a level and a content topic. I split them into 40 training tasks, 30 development tasks and 60 held-out tasks, with no topic shared between sets. Development tasks were for iterating. The held-out set was touched only for final checks. Confirmation runs used fresh tasks that neither version had ever seen.

Rules written in advance. Before each comparison I wrote down what would count as a win, in numbers. For the switch to a stronger writer with an editor gate (V5), the rule had three conditions: at least as many clean exercises, a flawed share no higher, and a cost per clean exercise no higher. All three had to hold.

Blind audits. Where the checks can’t see everything, a grader read the accepted exercises without knowing which version produced them. When both versions produced exactly the same exercise, word for word, it was graded once and the grade counted for both. The key that maps exercises to versions was opened only after every verdict was written down.

What moved the numbers

The first comparisons were about the pipeline’s shape. Same model (gpt-5.2), same 30 development tasks, four strategies.

Then the model. Same strategy, six model and reasoning settings, 30 tasks each.

On the 60 held-out tasks, the winning combination, which became V2, beat gpt-5.2 in the same pipeline 65% to 45%. On the tasks where exactly one of the two succeeded, it won 18 to 6, which is unlikely to be chance (two-sided sign test, p = 0.023). It was also cheaper per accepted exercise: $0.11 against $0.20.

Those two comparisons decided something important about what not to work on. I had planned to run an automated prompt optimiser, GEPA, over the generation and repair prompts. But strategy and model had moved acceptance by 40 to 60 points. A prompt optimiser could only tune wording, at about $50 a run, and it would only pay off at much larger volumes. I deferred it, and haven’t needed it since.

Ten days of decisions

Timeline from 27 September to 4 October: measure first; fix the judge and pick the strategy; pick the model by cost per accepted, which became V2; prompt optimiser deferred; coherence as its own stage adopted as V3; per-distractor gate rejected; design rules in the prompt adopted as V4; stronger writer with editor gate adopted as V5; three options not adopted; 30% mixed-grammar tolerance adopted as V6.
Every entry has a written rule and a result. Decisions get an ID in the project’s decision log (D21–D26 here), with the evidence linked. Each adopted change became a new pipeline version, V2 to V6.

Four of the ten entries are ideas that lost or were put on hold. I count that as the process working. In each case I wanted the idea to win, and without a rule written in advance I would have found a way to read the results generously.

IdeaRule written in advanceResultDecision
Coherence check inside the repair loopNo fewer accepted than a separate stage7 of 20 vs 8 of 20; it used up repair roundsSeparate stage: V3
Per-distractor gateAt least one more accepted, no higher cost2 of 10 vs 3 of 10, twice the costRejected
Design rules in the generation promptFour conditions on yield, flawed share and cost36 vs 32 of 60; flawed 12% vs 23% over 100 tasks eachAdopted as V4 (D24)
Stronger writer + editor gateClean ≥ baseline, flawed share ≤ baseline, cost per clean ≤ baseline30 vs 25 of 40; clean 21 vs 7; $0.098 vs $0.372 per cleanAdopted as V5 (D25)
Three options instead of fourCost per clean at least 15% lower$0.142 vs $0.137Not adopted
Pipeline versions: what V1–V6 mean
  1. V1
    Original pipelineuntil Sep 2026

    One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.

  2. V2
    Generate, verify, repair28 Sep 2026

    A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.

  3. V3
    + coherence check30 Sep 2026

    After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.

  4. V4
    + design rules in the prompt2 Oct 2026

    The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.

  5. V5
    + stronger writer, editor gate3 Oct 2026

    A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.

  6. V6
    + mixed-grammar tolerance4 Oct 2026

    Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.

Each version keeps everything before it. Charts and text name the version that produced a number.

If you make product decisions about AI behaviour

  1. Write the pass rule before the run, in numbers, and include cost.
  2. Keep a held-out set you don’t iterate on, and confirm on fresh tasks.
  3. Grade blind wherever humans or models judge quality.
  4. Test the big levers first, pipeline shape and model, before polishing prompts.
  5. Log what lost. It’s the most reused part of the record.
About the numbersSamples are small, 10 to 60 tasks per arm, so I treat single comparisons as directional and look for agreement across tests. Prices are OpenAI Batch API prices at the time. Model names are the ones available on my account in September and October 2026.