For months the site performed poorly in organic search. The frontend had started as a design in Lovable and grown with Claude Code, and along the way I had added prerendering, metadata, structured data and sitemaps. I asked a coding agent with SEO skills to find and fix problems. It found plausible improvements, some of them useful. The overall result didn’t move.
The real problem was simple: I’m not an SEO expert. I can work with product metrics, data, code and experiments, but technical SEO is a large set of interacting rules about crawling, rendering, canonical URLs, structured data, internal links and performance. I didn’t want the site’s quality to depend on me knowing which SEO question to ask next.
So the goal changed. Instead of using AI to help me do SEO, I built a system in which an agent acts as the SEO specialist and the developer: it inspects the live site, fixes problems, and proves that each fix worked.
The system
The first version was straightforward. n8n orchestrated: it picked checks from an SEO catalogue, tracked state and dispatched work. A Claude Code agent on GitHub investigated the repository, proposed fixes and implemented the approved ones. GitHub Actions ran tests, verification and production checks. Netlify provided preview deployments and production.
Editing n8n workflows by hand soon became the bottleneck, as retries, callbacks, observation windows and failure handling piled up. So I moved the control plane into GitHub. The workflows became version-controlled JSON, written and debugged by ChatGPT through its GitHub integration and deployed to n8n by GitHub Actions. That was the point where it stopped feeling like a collection of automations and started feeling like a system.
The hard part was verification
The plan was simple: take a check from the catalogue, let the agent investigate and fix, then verify. The verification turned out to be the weak link.
If a check required evidence from the live site, the agent might prove instead that the relevant code existed in the repository. If a check was defined on the rendered HTML, it might inspect the React source. If a performance check needed a specific Lighthouse signal, it might use a related PageSpeed number. Each substitution sounded reasonable. Each was a different check.
Coverage was a second problem. One representative page could pass while another page built from a different template stayed broken. It is the same lesson as with any LLM judge: an evaluator that can be renegotiated by the thing it evaluates isn’t an evaluator.
The frozen verifier
Now every fix comes with a verification contract. A proposal contains both the fix and the exact test that will prove it, and a separate gatekeeper reviews both. Then the verifier is frozen, before any implementation starts.
The baseline failure is the important step. A verifier is only useful if it can reproduce the defect it claims to test. The agent decides how to solve the problem. It can’t redefine afterwards what “solved” means. And if the right verifier doesn’t exist yet, the work goes on hold and building the verifier becomes a task of its own. It is never an excuse to use an easier check.
Three kinds of checks
The catalogue is split by how objectively a check can be decided. Deterministic checks have a machine-verifiable answer: does the canonical tag match, is the page in the sitemap, is the link present in the prerendered HTML? Deterministic-with-policy checks become mechanical once the site decides its own rules, for example whether the /exercise route should be crawlable at all. Semantic checks need judgement: does the page satisfy the search intent, is it meaningfully different from another page?
Ordinary problems, found systematically
None of the problems were exotic, which is partly the point. The grammar hub’s links to its categories existed only in client-side state, so the HTML Google received didn’t contain them. Exam-centre pages used JavaScript click handlers instead of real links. Some grammar content existed in the repository but was unreachable because two identifiers didn’t match. Unknown URLs returned a normal page instead of a not-found response. Structured data described content that wasn’t rendered. Each is easy to miss and cheap to fix. Together they were expensive.
Then search traffic moved
The first concentrated wave of these fixes was followed by a clear change in Google Search Console. Compared with the period just before it, organic clicks rose by about 77%. Impressions rose even faster and, once the initial spike settled, stayed about 223% above the earlier baseline. Later, impressions fell by about 44% as Google cut broad, low-click exposure, while clicks rose another 21% and click-through rate more than doubled.
This is an observation, not an experiment. There was no control group, and search traffic moves for many reasons. But the timing fits, the delay matches how long Google takes to recrawl, and the click growth held after the spike faded.
What comes next
Deterministic checks were the first stage. They can prove that a canonical tag is right. They can’t tell whether a page is genuinely useful. The next stage brings semantic checks, which need independent judges and stronger evidence, and above all a judge that isn’t the agent that made the change. The goal isn’t to make the agent more confident. It is to make the system better at knowing what it actually knows.
If you let an agent change production
- Agree on the test before the change, and freeze it.
- Require the test to fail first. A check that can’t see the defect proves nothing.
- Treat a missing verifier as work to do, not as permission to use an easier one.
- Separate the agent that changes things from the judge that evaluates them.
- Report outcomes honestly: observed patterns are useful, but they aren’t experiments.
