The old pipeline looked cheap. Generating an exercise cost about $0.036 including its own review, and an accepted one about $0.06. Then, in October, I re-checked the live library against the new standard, and only 16% of those exercises passed. They hadn’t been cheap at all. I had just been paying for them later: in learners meeting wrong answers, and now in replacements.
That is the whole argument of this article. The unit that matters is the cost of an exercise you would actually put in front of a learner.
Three numbers, one that counts
Every configuration I tested was compared on three numbers: cost per attempt, cost per accepted exercise and, once blind audits existed, cost per clean exercise, meaning accepted and graded as having nothing wrong with it. The first is what the invoice shows. The last is what you are buying.
The clearest case was the switch to a stronger writer with an editor gate (V5). A call to the new writer costs more than a call to the old one, and the editor gate adds another expensive call per exercise. Judged by the bill per call, I would have kept the old configuration.
Pipeline versions: what V1–V6 mean
- V1Original pipelineuntil Sep 2026
One model writes each exercise and reviews its own work. A draft that fails the review is rewritten from scratch until it passes or runs out of attempts. A separate checker pass flags problems afterwards.
- V2Generate, verify, repair28 Sep 2026
A mid-sized model (gpt-5.6-terra) writes. A new verifier checks every gap: free format rules, two small solvers that must agree on the one right answer, a judge for the German, the topic and the explanations, and an adjudicator when the solvers disagree. Only the gaps that fail are repaired, at most twice.
- V3+ coherence check30 Sep 2026
After every gap passes, a separate check reads the whole text: who does what, who a pronoun points to, whether facts contradict each other. A failure gets one text-only repair and a full re-check.
- V4+ design rules in the prompt2 Oct 2026
The rules for good wrong answers move from the checks into the writing prompt: each distractor is wrong in one way only, inside the grammar being practised, and no exercise can be solved by a surface pattern.
- V5+ stronger writer, editor gate3 Oct 2026
A frontier model (gpt-6.1-sol) writes and repairs. The same model then reads every finished exercise like a senior editor and rejects it for any of six named issue types: a second right answer, a double error, an option nobody would choose, a wrong key or explanation, a text problem, or mixed grammar.
- V6+ mixed-grammar tolerance4 Oct 2026
Up to 30% of an exercise’s gaps may touch a neighbouring grammar point, and a fixed phrase counts as one choice. A surface pattern across the whole exercise is still fatal. This is the version in production today.
Each version keeps everything before it. Charts and text name the version that produced a number.
The stronger writer passed more drafts at the first attempt, needed fewer repairs, and produced three times as many exercises that a grader found nothing wrong with: 21 clean out of 40 tasks instead of 7. Per clean exercise, it was 3.8 times cheaper. More expensive inputs, cheaper output.
Where the money goes
The October refill gave me the first full breakdown at production scale: 1,379 attempts, 974 accepted, every call attributed to the stage that made it.
Checking took 78% of the money. That surprised me at first, and then it didn’t. An exercise is written once. It is then read by two solvers, a judge, sometimes an adjudicator, a coherence check and the editor gate, and every repair sends it through the gap checks again. For 1,379 attempts that came to more than 14,000 model calls.
| Stage | Model | Calls | Cost |
|---|---|---|---|
| Solvers (two per check) | gpt-5-mini | 5,064 | $23.80 |
| Judge | gpt-5.2 | 2,532 | $17.98 |
| Editor gate | gpt-6.1-sol | 1,603 | $12.61 |
| Writing first drafts | gpt-6.1-sol | 1,379 | $9.07 |
| Adjudicator (solvers disagree) | gpt-5.2 | 437 | $3.62 |
| Repairs after the editor gate | gpt-6.1-sol | 491 | $3.44 |
| Gap repairs | gpt-6.1-sol | 540 | $3.43 |
| Coherence check | gpt-5.2 | 1,865 | $2.75 |
| Strict re-check of repaired gaps | gpt-5.2 | 317 | $1.23 |
| Coherence repairs | gpt-6.1-sol | 193 | $1.11 |
| Total | 14,421 | $79.03 |
I think this ratio is right, and I’d defend it to anyone who asks why checking costs almost seven times as much as writing. In a product where a wrong exercise does real damage, verification is the product, and writing is the cheap part. The question is how to spend less on it without seeing less.
Two answers so far. The solvers, the biggest line, already use the smallest model that measured well for the job. And better first drafts mean fewer repairs and fewer re-checks. That is why the distractor rules moved from the checker into the writing prompt: first drafts passing every gap check went from 8% to 32% on the same tasks, and every draft that passes first time skips a repair and a whole round of checks.
Batch prices, batch time
Everything runs through OpenAI’s Batch API at half the normal price. The cost is time. A run is a sequence of rounds (write, check, repair, check again), and each round waits for its slowest batch. In the October refill, batches for the small models finished in minutes, while batches for the large model took anywhere from two to fourteen hours. The 1,379 attempts needed 21 rounds over four days.
For a library refill that is fine. For anything a learner is waiting for it is not, and the two should never share a mode. The improvement I’d make next is cheap: once only a few dozen requests remain, send the tail through the normal API instead of the Batch API. That would cost a few dollars and save a day or two.
Making long runs safe to leave alone
A run that lasts four days has to be safe to walk away from.
- A budget cap per run. The run stops cleanly before it would cross the cap, and resumes only if the cap is raised.
- Quota errors stop the run, they don’t fail items. The account hit its spending limit twice during the refill. Each time the run saved its state, printed how to resume and exited. No exercise was ever marked as failed because of an API error.
- Everything resumes from cached responses. Running the same command again collects finished batches and carries on.
Compute that isn’t a model call
Not every cost is an API call. The similarity analysis that keeps exercise sequences varied (described here) is plain CPU work. It runs on a rented CPU machine on Vast.ai, picked from offers under $0.10 an hour. The machine gets the code and a state file, never the database credentials, and only my own machine writes results back. A rented server can’t touch production data even if something goes wrong on it.
If you’re budgeting an LLM feature
- Compare configurations by cost per good output. Price per token routinely points the wrong way.
- Expect checking to cost more than generating, and treat that as the price of being right.
- Spend less on checks by improving first drafts, not by removing checks.
- Keep batch work and interactive work apart, and push long batch tails through the normal API.
- Make long runs boring: budget caps, quota stops, resumable state.
