Engineering

How we verify AI-written pull requests before they merge: our evidence gate

Associates AI ·

AI-written code is usually almost right, which is the expensive kind of wrong. Here is the evidence gate every change passes in our own repositories, two real examples of what it caught, what it costs, and how to build the same gate yourself.

How we verify AI-written pull requests before they merge: our evidence gate

Why AI-written pull requests need a different kind of check

The failure mode of AI-written code is rarely a crash. It is a change that looks right, passes a skim, and is wrong in one place nobody looked.

The Stack Overflow 2025 Developer Survey puts numbers on the feeling. More developers actively distrust the accuracy of AI tools (46%) than trust it (33%), and the biggest single frustration, cited by 66% of developers, is dealing with "AI solutions that are almost right, but not quite". 45.2% cited debugging AI-generated code as more time-consuming.

"Almost right" is hard to catch by reading, because the diff reads fine. So we stopped asking reviewers to be more careful and started asking every pull request to bring evidence. This post is the gate we run on our own repositories: what it requires, what it caught recently, what it costs, and where it still leaks.

What our gate requires on every pull request

The gate assumes the author, often an AI agent, can be confidently wrong. We ask every pull request for these:

  1. A spec with numbered acceptance criteria. The tracked issue holds the approved spec. The code has to prove it, not replace it. If you cannot say what "done" means, you cannot check it.
  2. A test seen failing, then passing. A new assertion must be observed failing against the broken state before the fix goes in. A test that has never failed has not shown it can fail.
  3. A fresh-context review. A reviewer that never saw the authoring conversation reads the whole diff. It looks for security regressions, assertions that pass while proving nothing, anything that breaks boot or deploy, and complexity the spec did not ask for. Every finding gets fixed, not just the blocking ones.
  4. Green checks on the head commit. Not on an earlier commit, and not assumed from a branch rule. Someone reads the actual check result for the exact commit being merged.
  5. Evidence, blast radius and rollback in the body. The pull request description pastes real command output, says what breaks if the change is wrong and for whom, and states the exact way back.
  6. A clean merge state. No unresolved review threads, and a branch that is current with main.

Changes that reach production infrastructure add a second layer. They run on a staging instance of the real platform first, exercising the checklist rows the change can touch. A platform version upgrade has to pass on a freshly built instance and on an existing instance upgraded from the old version. After merge, one canary instance takes the change alone. A read-only verification suite reports whether it booted cleanly, whether the gateway is healthy, whether channels started, whether there are any authentication errors, and whether one real agent turn completed. Only then does the rest of the fleet move, and we never run two production changes at once.

That verification suite exists because of two incidents. In one, an agent answered nothing for days while the health metric stayed green, because the metric never made a real model call. In the other, a bad image boot-looped every instance because nobody had booted it before promoting it. Both passed every signal the old gate checked.

What the gate caught, and what it missed: two real examples

These are lightly redacted and described by type.

A deploy tool that failed far from its cause. This week a rollout failed twice, forward and rollback, with a cryptic syntax error. The cause was a build step that downloaded a deploy tool with no check for failure. A transient non-success response was saved as if it were the program. The same URL worked seconds later, and the instance was never touched, because the failure came before that step.

The fix was small: fail on HTTP errors, retry, check the downloaded file's type, and stop with a clear message. The gate still did its work on a small change:

  • The new test failed against the old behavior (four of six checks) and passed after (six of six).
  • The fresh-context review raised two important findings in its first round. One was a silent swallow of an archive-extraction failure. Both were fixed before merge.
  • The pull request named its blast radius: every deploy and destroy runs through that step, so a false-positive check would block all of them. That is why a real staging deploy was required, not just a unit test.
  • After the production apply, the first real production run logged the new download step succeeding, and the downstream pipeline finished green.

A fix that did not do what it claimed. We want thinking turned on by default for every model that supports it. A merged change widened support for two models so it would work. Measuring the real runtime in the follow-up work, not reading the diff, showed both still resolved to thinking off. The diff only widened a list, so reading it could not have shown this. That is a miss: the change passed its gate and merged, and the measurement that caught it came after.

The follow-up spec gained an acceptance criterion, and the fix recorded the failing output before the change and the passing output after: both models went from off to medium. It also added a check to the image build that fails the build if any model that should reason resolves to off, across the whole model catalog. A check that fails the build is cheaper than a silent setting in a customer's agent.

What the gate costs in time

The record shows uneven numbers, not a clean figure. The download fix went from issue to merged pull request in about twenty hours, but only about twelve minutes passed between opening the pull request and merging it, because the review and the failing-then-passing test had already happened before it opened. The staging deploy came after merge, on the way to production, not before the pull request. The cost sits upstream of the pull request, in writing the spec, watching the test fail, and a review round or two.

Infrastructure changes cost more because staging and the canary are real instances on a real clock, so we batch everything that needs staging into one window. We have not measured cost per change, so we will not claim a number.

What we still get wrong

  • The gate is not coverage. A test seen failing proves the test can fail. It does not prove the suite covers what matters. The download fix checks the file type, not a published checksum, because none was pinned.
  • Some evidence arrives after merge. For infrastructure changes, the staging run and the first real production run can both come after the merge, and the production run is the strongest proof. We say so in the pull request instead of pretending staging covered it, but it is a gap.
  • Review happens before merge, not always before the expensive step. A reviewer reads a diff. A build, a bake or a rollout can already have cost time by the time something is found.
  • Small fixes skip some ceremony. A small bug fix can merge against a clear bug report rather than numbered criteria, and without a separate approving review once the fresh-context review is clean.
  • A passing gate can still be wrong. The incidents above passed the gate that existed at the time, and so did the thinking change. We close each gap when we find it, which means the next one is out there.

How to install this gate on your own repository

You do not need our platform.

  1. A pull request template with three required sections: evidence (pasted output, not "tested locally"), blast radius, and rollback. A template makes the empty section visible.
  2. Branch protection that requires your CI to pass on the head commit and at least one approving review, blocks merging with unresolved threads, and requires the branch to be up to date.
  3. A fresh-context reviewer. A person, or a second AI model session, that never saw the authoring conversation and reads only the diff and the spec. Tell it what to hunt for: tests that cannot fail, anything that breaks boot or deploy, scope creep.
  4. A failing-test-first rule. Require the pull request to link the run where the new test failed against the old code.
  5. Staging and a canary for infrastructure. Run the change on a staging environment that matches production, then on one production instance alone, and read real signals before moving the rest. A green health check is the start of the evidence, not the end.
  6. Write down what remains uncertain. Ask the author to say what the evidence does not cover. The answer is where your next test comes from.

If you want a second reviewer on your pull requests

Associates AI is the managed OpenClaw hosting platform, and we run this gate on our own engineering first. Our offer is simple: connect a repo and get a second review of the PRs you already have, before anything writes code: findings on every PR it reviews, and test evidence on every PR the Engineer Teammate opens. In our own use, the reviewer surfaced useful findings the existing review missed on about half the PRs it reviewed; the measurement window is running. If you want to talk it through, get in touch through the site.

MH

Written by

Mike Harrison

Founder, Associates AI

Mike is a self-taught technologist who has spent his career proving that unconventional thinking produces the most powerful solutions. He built Associates AI on the belief that every business — regardless of size — deserves AI that actually works for them: custom-built, fully managed, and getting smarter over time. When he's not building agent systems, he's finding the outside-of-the-box answer to problems that have existed for generations.

More from the blog

Ready to put AI to work for your business?

Get started today. Hire your first Teammate in minutes and put it to work on what you're reading about.

Get Started