On 5 October I ran every live exercise on InfiniteGrammar.de through the same checks a new exercise has to pass in pipeline V6. I expected a lot of failures. The library had been written by an older pipeline that graded its own work, and a partial clean-up a week earlier had already taken down more than 500 exercises, most of them for ambiguous answers.
I didn’t expect 16%.
133 of 829 passed. The other 696 failed somewhere, and more than half of those had a distractor that was also correct, the same problem that started the whole quality project.
Was the gate too strict?
251 of the failures came from the final check, the editor gate. I knew from calibration that it’s strict. On test exercises I had graded blind, only about one in four of its rejections was an exercise graded flawed. Most of the rest were usable, often with a minor issue.
So taking down everything that failed meant taking down some good exercises. I decided to do it anyway, for two reasons. Learners can’t see which pipeline wrote an exercise, so one standard for everything they see is the only promise I can actually keep. And replacements were cheap: about eight cents each, and they pass the same bar. The takedown wrote an undo file, so it was reversible. A learner who stops trusting the answers isn’t.
Pipeline versions: what V1–V6 mean
- V1Original pipelineuntil Sep 2026
One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.
- V2Generate, verify, repair28 Sep 2026
A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.
- V3+ coherence check30 Sep 2026
After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.
- V4+ design rules in the prompt2 Oct 2026
The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.
- V5+ stronger writer, editor gate3 Oct 2026
A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.
- V6+ mixed-grammar tolerance4 Oct 2026
Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.
Each version keeps everything before it. Charts and text name the version that produced a number.
How much to generate
A flat target of, say, 200 exercises per section would have been simple and wrong. Some sections had many active learners working through them quickly. Others had just a few keen learners. But the product promises “infinite grammar exercises in every grammar section, for every learner”. So each of the 86 sections got its own target: the number of exercises its most active learners used in their first month there.
That came to 808 new exercises needed. At the yields measured earlier for each level, between 50% and 80%, that meant 1,379 attempts, estimated at $76.
The order of operations
The obvious sequence is to take the failures down and then refill. It would have left 55 sections completely empty for the four days the refill took. So the sequence went the other way.
The other constraint was learner progress. Learners work through a section in order, and the app serves the lowest active exercise number above the last one they completed. Taken-down exercises keep their numbers and are simply skipped. New exercises are added after the end. So nobody repeats an exercise and nobody skips one they haven’t seen. Before the real run, I did a dry run of both the takedown and the refill, to make sure the whole operation would go through without a flaw.
What came out
The refill made 1,379 attempts over 21 batch rounds and accepted 974, or 71%, for $78.14. A random sample of 20 read as 14 clean, 6 with minor issues and none flawed. The live library went from 829 exercises of mixed quality to 1,107 that all pass today’s checks.
If you’re cleaning up generated content
- Hold everything users see to one standard, and say openly what it costs you.
- Generate the replacements before you take anything down.
- Size content by your most engaged users, not by averages.
- Keep bulk changes reversible: dry runs, undo files, counts before and after.
