The OpenClaw Memory Plugin Compromise: Why a Trusted Skill Needs Its Own Boundary
A compromised memory integration ran a credential-hunting payload during ordinary OpenClaw use. The...
AI-written code is usually almost right, which is the expensive kind of wrong. Here is the evidence gate every change passes in our own repositories, two real examples of what it caught, what it costs, and how to build the same gate yourself.
The failure mode of AI-written code is rarely a crash. It is a change that looks right, passes a skim, and is wrong in one place nobody looked.
The Stack Overflow 2025 Developer Survey puts numbers on the feeling. More developers actively distrust the accuracy of AI tools (46%) than trust it (33%), and the biggest single frustration, cited by 66% of developers, is dealing with "AI solutions that are almost right, but not quite". 45.2% cited debugging AI-generated code as more time-consuming.
"Almost right" is hard to catch by reading, because the diff reads fine. So we stopped asking reviewers to be more careful and started asking every pull request to bring evidence. This post is the gate we run on our own repositories: what it requires, what it caught recently, what it costs, and where it still leaks.
The gate assumes the author, often an AI agent, can be confidently wrong. We ask every pull request for these:
Changes that reach production infrastructure add a second layer. They run on a staging instance of the real platform first, exercising the checklist rows the change can touch. A platform version upgrade has to pass on a freshly built instance and on an existing instance upgraded from the old version. After merge, one canary instance takes the change alone. A read-only verification suite reports whether it booted cleanly, whether the gateway is healthy, whether channels started, whether there are any authentication errors, and whether one real agent turn completed. Only then does the rest of the fleet move, and we never run two production changes at once.
That verification suite exists because of two incidents. In one, an agent answered nothing for days while the health metric stayed green, because the metric never made a real model call. In the other, a bad image boot-looped every instance because nobody had booted it before promoting it. Both passed every signal the old gate checked.
These are lightly redacted and described by type.
A deploy tool that failed far from its cause. This week a rollout failed twice, forward and rollback, with a cryptic syntax error. The cause was a build step that downloaded a deploy tool with no check for failure. A transient non-success response was saved as if it were the program. The same URL worked seconds later, and the instance was never touched, because the failure came before that step.
The fix was small: fail on HTTP errors, retry, check the downloaded file's type, and stop with a clear message. The gate still did its work on a small change:
A fix that did not do what it claimed. We want thinking turned on by default for every model that supports it. A merged change widened support for two models so it would work. Measuring the real runtime in the follow-up work, not reading the diff, showed both still resolved to thinking off. The diff only widened a list, so reading it could not have shown this. That is a miss: the change passed its gate and merged, and the measurement that caught it came after.
The follow-up spec gained an acceptance criterion, and the fix recorded the failing output before the change and the passing output after: both models went from off to medium. It also added a check to the image build that fails the build if any model that should reason resolves to off, across the whole model catalog. A check that fails the build is cheaper than a silent setting in a customer's agent.
The record shows uneven numbers, not a clean figure. The download fix went from issue to merged pull request in about twenty hours, but only about twelve minutes passed between opening the pull request and merging it, because the review and the failing-then-passing test had already happened before it opened. The staging deploy came after merge, on the way to production, not before the pull request. The cost sits upstream of the pull request, in writing the spec, watching the test fail, and a review round or two.
Infrastructure changes cost more because staging and the canary are real instances on a real clock, so we batch everything that needs staging into one window. We have not measured cost per change, so we will not claim a number.
You do not need our platform.
Associates AI is the managed OpenClaw hosting platform, and we run this gate on our own engineering first. Our offer is simple: connect a repo and get a second review of the PRs you already have, before anything writes code: findings on every PR it reviews, and test evidence on every PR the Engineer Teammate opens. In our own use, the reviewer surfaced useful findings the existing review missed on about half the PRs it reviewed; the measurement window is running. If you want to talk it through, get in touch through the site.
Written by
Founder, Associates AI
Mike is a self-taught technologist who has spent his career proving that unconventional thinking produces the most powerful solutions. He built Associates AI on the belief that every business — regardless of size — deserves AI that actually works for them: custom-built, fully managed, and getting smarter over time. When he's not building agent systems, he's finding the outside-of-the-box answer to problems that have existed for generations.
More from the blog
A compromised memory integration ran a credential-hunting payload during ordinary OpenClaw use. The...
The new Blueprint Alliance treats AI coworkers as identities that must be discovered, owned, scoped,...
Jobber's new Teammate does three things most AI products still avoid: briefs the owner, prepares wor...
Want to go deeper?
Get started today. Hire your first Teammate in minutes and put it to work on what you're reading about.
Get Started