This week the agent and I designed the database side of the admin dashboards. It wrote the migrations, tested them against a throwaway Postgres loaded with an export of the data, caught a wrong assertion in its own tests, fixed it, and handed me the commands to run against production, with the output I should expect from each. I read them, ran them and pasted the output back. The numbers matched.
That split is the whole operating model. The agent does the engineering. I own the direction and the risk.
Shared memory: the roadmap file
An agent forgets everything between sessions, and long sessions get summarised and lose detail. So the project has one file that every session reads first: a roadmap with the goal, the task tables, the results of every run and a table headed “Decisions already made (don’t re-discuss)”. It holds numbered decisions, D1 to D27 so far, each with a one-line rationale. Some are big (which pipeline is in production). Some are small and save a surprising amount of time (no human labelling, OpenAI only for generation, German only).
The rule for the agent is simple: when a task is finished, update its status and append the result to the log in the same pull request. The rule for me is just as simple: if I change my mind, I add a decision rather than arguing in chat. That one heading, “don’t re-discuss”, has saved me more time than any prompt I have written.
Budgets, in money and in blast radius
Every working session that spends API money has a budget. Test runs stay small, 50 exercises or fewer unless a task says otherwise, and long batch runs go to my own machine rather than a cloud session. The pipeline enforces its own caps too: a run stops cleanly before crossing its budget and can be resumed.
The other budget is blast radius. Cloud sessions for content generation run against a local copy of the database loaded from a fixture file. Unlike the sessions working on the frontend, they don’t connect to production at this stage. Every write to production data is run by me, from a short runbook the agent writes: the command, a dry run first, the output I should expect, and how to undo it. Once I’m satisfied with the quality of the generation pipeline, I will extend its permissions.
Gates for code, not just for content
Changing a prompt or a model setting in this project changes what learners see, so those changes have their own regression test. It runs the production configuration on 20 fixed tasks and on a set of known tricky texts, costs about $2, and fails if the pass rate drops by more than five points or the cost rises by more than 25%. Any change to the generator, the prompts, the pipeline code or the verifier has to pass it before merging. On top of that there are more than 260 unit tests, one feature branch and pull request per task, and nothing pushed straight to main.
API errors get special treatment, because they are easy to misread. A quota error must stop the run, save state and print how to resume. A rate limit must back off and retry. And no API error may ever be recorded as a quality verdict. An exercise that failed because the API was down is not a failed exercise.
How I review
I don’t read every line of code. I read the summary in plain language, the numbers, and the full diff of anything that touches data or decides what learners see. I ask for a recommendation, not a menu of options. When I disagree, I say so and the decision goes into the log. Most of my time goes into three things: deciding, reading output (the exercises themselves, not reports about them) and saying no.
Three mistakes and the rules they produced
A commit that carried someone else’s draft. Another session, working on the frontend, had drafted a different definition of “verified” directly into one of the spec files and left it uncommitted for review. A commit in my session edited the same file and swept the draft in with it. The repository then held two documents that contradicted each other on the most important number in the dashboard, which was already showing 0% of exercises verified instead of 100%. The fix took minutes. The rule: when using parallel sessions, have the diffs reviewed before committing, and leave other sessions’ edits out.
A long run tied to the editor. The four-day refill at first ran as a background process inside the editor session. Closing the editor stopped it in the middle of the run. Nothing was lost, because every run resumes from cached results, but until I noticed, no new batches went out. The rule: long runs either run detached or are cheap to restart, and the restart command is always written down.
Acting on a check before calibrating it. This one was mine. The agent built a first verifier quickly and I used it at once to take 411 exercises offline. Later calibration showed it raised false alarms 5–8% of the time. The rule: a check earns the right to act only after it has been tested on data with known answers (how I do that now).
What this changes about product work
Less than people expect. Deciding what to build, specifying it precisely, checking the result against the specification and saying no to the rest is still the job. What changes is the loop. A specification turns into working, tested code in hours, so the bottleneck moves to the quality of the decisions and the reading. That is exactly where a product manager’s time has to go, with a human team or an AI one.
If you’re working with a coding agent
- Keep one file of decisions and results that every session reads first.
- Give every session a budget, in money and in what it’s allowed to touch.
- Keep production writes for yourself, with a runbook, a dry run and an undo.
- Put a regression test in front of anything that changes user-facing behaviour.
- Turn every mistake into a written rule the agent reads next time.
