For months the site performed poorly in organic search. The frontend had started as a design in Lovable and grown with Claude Code, and along the way I had added prerendering, metadata, structured data and sitemaps. I asked a coding agent with SEO skills to find and fix problems. It found plausible improvements, some of them useful. The overall result didn’t move.

The real problem was simple: I’m not an SEO expert. I can work with product metrics, data, code and experiments, but technical SEO is a large set of interacting rules about crawling, rendering, canonical URLs, structured data, internal links and performance. I didn’t want the site’s quality to depend on me knowing which SEO question to ask next.

So the goal changed. Instead of using AI to help me do SEO, I built a system in which an agent acts as the SEO specialist and the developer: it inspects the live site, fixes problems, and proves that each fix worked.

The system

The first version was straightforward. n8n orchestrated: it picked checks from an SEO catalogue, tracked state and dispatched work. A Claude Code agent on GitHub investigated the repository, proposed fixes and implemented the approved ones. GitHub Actions ran tests, verification and production checks. Netlify provided preview deployments and production.

Orchestration stack: SEO check catalogue, n8n for orchestration and state, GitHub and Claude to investigate, propose and fix, GitHub Actions to test and verify, Netlify for preview and production.
The orchestration stack. Every issue, branch, pull request and result is visible in GitHub.

Editing n8n workflows by hand soon became the bottleneck, as retries, callbacks, observation windows and failure handling piled up. So I moved the control plane into GitHub. The workflows became version-controlled JSON, written and debugged by ChatGPT through its GitHub integration and deployed to n8n by GitHub Actions. That was the point where it stopped feeling like a collection of automations and started feeling like a system.

The hard part was verification

The plan was simple: take a check from the catalogue, let the agent investigate and fix, then verify. The verification turned out to be the weak link.

If a check required evidence from the live site, the agent might prove instead that the relevant code existed in the repository. If a check was defined on the rendered HTML, it might inspect the React source. If a performance check needed a specific Lighthouse signal, it might use a related PageSpeed number. Each substitution sounded reasonable. Each was a different check.

Why verification drifted: the original check asks whether production shows X; the agent implements a fix; the exact check is inconvenient; the agent runs a related check Y; Y passes; the issue is reported as fixed.
Verification drift. The agent was, in effect, allowed to change the test after seeing its own implementation.

Coverage was a second problem. One representative page could pass while another page built from a different template stayed broken. It is the same lesson as with any LLM judge: an evaluator that can be renegotiated by the thing it evaluates isn’t an evaluator.

The frozen verifier

Now every fix comes with a verification contract. A proposal contains both the fix and the exact test that will prove it, and a separate gatekeeper reviews both. Then the verifier is frozen, before any implementation starts.

The frozen verifier: catalogue check, assessment finds an issue, proposal with fix and verification contract, gatekeeper, freeze verifier (current production must fail), implementation, deploy preview must pass, production must pass the same verifier.
The verifier is frozen before implementation. It has to fail on production first, which proves it can see the defect at all.

The baseline failure is the important step. A verifier is only useful if it can reproduce the defect it claims to test. The agent decides how to solve the problem. It can’t redefine afterwards what “solved” means. And if the right verifier doesn’t exist yet, the work goes on hold and building the verifier becomes a task of its own. It is never an excuse to use an easier check.

Three kinds of checks

The catalogue is split by how objectively a check can be decided. Deterministic checks have a machine-verifiable answer: does the canonical tag match, is the page in the sitemap, is the link present in the prerendered HTML? Deterministic-with-policy checks become mechanical once the site decides its own rules, for example whether the /exercise route should be crawlable at all. Semantic checks need judgement: does the page satisfy the search intent, is it meaningfully different from another page?

Three types of SEO checks: deterministic, deterministic with policy, and semantic.
Three kinds of checks. Much of the work turned out to be making implicit product decisions explicit enough for software to test.

Ordinary problems, found systematically

None of the problems were exotic, which is partly the point. The grammar hub’s links to its categories existed only in client-side state, so the HTML Google received didn’t contain them. Exam-centre pages used JavaScript click handlers instead of real links. Some grammar content existed in the repository but was unreachable because two identifiers didn’t match. Unknown URLs returned a normal page instead of a not-found response. Structured data described content that wasn’t rendered. Each is easy to miss and cheap to fix. Together they were expensive.

Then search traffic moved

The first concentrated wave of these fixes was followed by a clear change in Google Search Console. Compared with the period just before it, organic clicks rose by about 77%. Impressions rose even faster and, once the initial spike settled, stayed about 223% above the earlier baseline. Later, impressions fell by about 44% as Google cut broad, low-click exposure, while clicks rose another 21% and click-through rate more than doubled.

Observed outcomes after the first remediation wave: clicks from an index of 100 to 177; impressions from 100 to 323; later impressions down 44% while clicks rose another 21% and CTR more than doubled.
Observed, not proven. The improvement started after the fixes, arrived with a plausible recrawl delay and survived the end of the impression spike.

This is an observation, not an experiment. There was no control group, and search traffic moves for many reasons. But the timing fits, the delay matches how long Google takes to recrawl, and the click growth held after the spike faded.

What comes next

Deterministic checks were the first stage. They can prove that a canonical tag is right. They can’t tell whether a page is genuinely useful. The next stage brings semantic checks, which need independent judges and stronger evidence, and above all a judge that isn’t the agent that made the change. The goal isn’t to make the agent more confident. It is to make the system better at knowing what it actually knows.

If you let an agent change production

  1. Agree on the test before the change, and freeze it.
  2. Require the test to fail first. A check that can’t see the defect proves nothing.
  3. Treat a missing verifier as work to do, not as permission to use an easier one.
  4. Separate the agent that changes things from the judge that evaluates them.
  5. Report outcomes honestly: observed patterns are useful, but they aren’t experiments.