How we run our own engineering with agents

We build Associates AI with coding agents that follow the same trail our engineering teammates use. Every change to the platform goes through it: an issue with numbered acceptance criteria, a test that fails before the change and passes after, a review the engineer has to answer, a stated blast radius and rollback, and a merge only once CI is green and every finding is answered. Below are four real pull requests from our private platform repository, lightly redacted.

A public demo repository where our product, engineer and QA teammates work in the open is opening soon.

A rotated webhook secret takes effect within a minute

The problem

The inbound webhook handler cached each server's signing secret with no time limit. After a rotation, the old secret kept working and the new one was rejected for an unbounded time.

Acceptance criteria

  • On a signature mismatch, drop the cached secret, fetch it once more and re-check before rejecting.
  • Cached secrets expire after 60 seconds, so a rotated-out secret stops working even without a mismatch.
  • A wrong secret costs exactly one extra fetch, and a failed fetch fails closed.

Failing, then passing

The rotation tests were run against the old handler first and failed: a rejection where an accept was expected, and one fetch where two were expected. After the change the next request after a rotation accepts the new secret and rejects the old one.

Review findings

  • Fresh-context review: nothing bounded how long a revoked secret stayed valid. We added the 60-second expiry and a test for it.
  • It also spotted the same cache shape in the webhook settings, which got its own pull request.

Blast radius and rollback

Blast radius: the webhook handler only, at most one extra secret read per failed signature and one per server per minute. Rollback: revert the commit; the cache lives in memory, so there is nothing to unwind.

Merge

CI green on the final commit, merged once every review finding was answered, then proven on a staging server by rotating a secret and checking the old one is refused.

A flaky tool download can no longer break a deploy

The problem

Our deploy tooling downloaded two tools without checking the response. A transient error page was saved as the tool, and the deploy failed minutes later with a confusing error.

Acceptance criteria

  • Downloads fail on HTTP errors and retry.
  • The downloaded file must be a real archive or executable, and a failed unpack fails the step.
  • The failure message says the download failed, and the operator runbook lists the symptom.

Failing, then passing

A new shell test failed four of six checks against the old download step and passes all six after: a failed response must never leave an executable behind.

Review findings

  • Fresh-context review found that a failed unpack was silently ignored. Fixed in the same pull request.

Blast radius and rollback

Blast radius: the install step of every deploy, so a false positive would block all deploys. That is why it ran on staging first. Rollback: revert.

Merge

CI green, merged once every review finding was answered, then proven on staging.

A health check stops reporting a usage-limit pause as a login failure

The problem

Our fleet verification script counted a harmless log line, written after a model provider's rate limit, as an authentication failure. A healthy server showed a red result.

Acceptance criteria

  • That line no longer counts as an authentication failure; real authentication failures still fail the check.
  • A separate warning row shows when usage limits are hit, so the pressure stays visible.
  • The verification runbook is updated.

Failing, then passing

A new fixture with only the harmless line failed against the old check and passes after. Putting the old marker back makes the marker-list test fail.

Review findings

  • The automated QA reviewer found that a permission error on one lookup was reported as a generic error and the check carried on, unlike its siblings. Fixed to report the denial and stop, with a test that denies only that call.

Blast radius and rollback

Blast radius: operator tooling only, with no change to servers or infrastructure. Rollback: revert.

Merge

CI green, merged once every review finding was answered.

Memories an agent deliberately saves are always searchable

The problem

A memory an agent was explicitly asked to save could land in a hidden tier and never come back in a search. Two search paths also kept separate filters that agreed only by coincidence.

Acceptance criteria

  • Every explicit save lands in the searchable tier.
  • Both search paths share one filter and return the same results for the same data.
  • A one-time, repeatable backfill promotes earlier explicit saves.

Failing, then passing

Each new assertion was seen failing against the old code first. A three-row fixture proves both search paths return the same memories, and a guard fails if a hand-copied filter comes back.

Review findings

  • Fresh-context review: two important findings and four nits, all fixed, including a fixture row that did not actually exercise per-user visibility.
  • A second automated reviewer said the guard was too literal. Fixed in a follow-up commit.

Blast radius and rollback

Blast radius: which memories every agent can find, while ranking is unchanged. Rollback: revert; promoted rows stay promoted, so no data is lost.

Merge

CI green, merged once every review finding was answered, proven on staging before the release.

Want this on your repository?

Book a 20-minute look