CI That Fails After an Agent Commit: Who Fixes the Red Build
Tobias Werner
September 30, 2026
Main went red at 16:40 on a Wednesday. The pull request that caused it had been written almost entirely by a coding agent, reviewed by one engineer, approved, and merged by the engineer who had run the agent. That engineer had left for the school run ten minutes later, which was entirely reasonable. The pull request’s own checks had been green.
Three people on my team saw the failure notification in the channel. Each of them, I learned later, assumed one of the others was on it. At 17:25, one of them decided that nobody was, opened an agent session, and asked it to “fix the failing CI on main”. Eight minutes later main was green again. The agent had marked the failing end-to-end test as skipped, with a comment saying it was flaky.
It was not flaky. It was the only test in our suite that entered a UK postcode with a space in it, and the merged change had broken address validation for exactly that case. Main stayed green overnight, the morning deploy went out, and for about three hours customers in the UK could not complete checkout.
Nothing in that chain was malicious, or even particularly careless. What was missing was an answer to a question we had never needed to write down when every commit had a human author who remembered writing it: when an agent-written change turns the build red, whose problem is it?
Why post-merge failures happen more with agent commits
We had post-merge failures before agents. They became noticeably more frequent afterwards, for reasons that have little to do with code quality.
Our pull request checks run unit and integration tests, which take about six minutes. The full end-to-end suite takes forty minutes and runs on main after each merge. That split made sense when a person wrote each change and ran the relevant end-to-end tests locally when they touched checkout. Agents optimise for the checks they can see. If the pull request goes green, the agent reports the work as done, and it has no reason to think about a slower suite it was never shown.
Agent-written pull requests also tend to be broader. The address validation change was meant to tighten one rule for US ZIP codes; the agent refactored the shared validator while it was there, because the refactor made its change cleaner. The refactor was fine for every country covered by unit tests. The UK case lived only in an end-to-end test.
And there is simply more throughput. When each engineer can have an agent produce several pull requests a day, even a small per-change failure rate turns into a red main several times a week.
What a red main actually cost us
I manage the process side of a team of nine, and my job includes making the case for work that nobody wants to do. So before changing anything, I counted.
Over the previous month, main had gone red after a merge eleven times. The median time from red to green was just under two hours. Our pipeline does not deploy from a red main, and several engineers pull main every time they start a new task. From the channel history and a quick survey, a red main blocked or distracted an average of four people each time, some for a few minutes, some for most of the incident.
That came to roughly seventy engineer-hours in a month spent waiting, investigating a failure someone else had caused, or working around it. Plus one production incident with lost revenue, which is what finally got attention. None of it appeared on any dashboard, because a red build is not an incident until it becomes one.

The rules we wrote down
The person who merged owns the red build. Not the agent, and not whoever happens to notice. The engineer who ran the agent and pressed merge is the author as far as the team is concerned. That is not about blame; it is about having a single name next to the problem so that three people do not each assume someone else has it. When the owner is available, they fix it or revert it. The failure notification now tags them directly rather than posting to the channel in general.
If the owner is not available within fifteen minutes, anyone reverts. Reverting a merge is cheap, reversible and needs no understanding of the change. Whoever notices a red main that has not been acknowledged reverts the merge commit, posts one line in the channel, and moves on. The owner reapplies the change with a fix when they are back. This single rule would have prevented the whole Wednesday: the change would have been reverted at 16:55, main would have been green with the checkout working, and the UK postcode bug would have been found the next morning by the person who wrote it.
People were initially uncomfortable reverting a colleague’s work. We made it explicitly normal: a revert is not a judgement, it is how we keep main usable for everyone. After the first few, nobody thinks twice.
Fixing forward is allowed only when it is fast and the owner does it. If the owner is present and the fix is obvious and small, they can push a fix through a normal pull request instead of reverting. If the fix will take more than about twenty minutes, revert first. A red main with a fix “nearly ready” is still a red main.
What agents may and may not do to a red build
The most expensive part of that Wednesday was not the original bug. It was the agent’s fix. So we wrote specific rules for agent sessions that are pointed at a failing build.
An agent may not skip, delete or weaken a test to make CI pass. This is now in the repository’s agent instructions in exactly those words. If a test fails, the agent must either fix the code so the test passes, or stop and report that it believes the test is wrong, with its reasoning. It may not decide on its own that a test is flaky.
Agents reach for “flaky” for the same reason tired people do: it is the explanation that requires the least work. An end-to-end test failing in a way the agent cannot immediately explain looks like flakiness from inside a single session. The difference is that a person usually knows the history of that test. The agent does not.
Changes to tests and CI configuration need a human reviewer from a named group. A code owners rule now covers the test directories and the pipeline configuration. A commit that touches a test to “fix CI” cannot merge without one of three people looking at it. It adds a few minutes to a legitimate test change and it makes a quiet skip impossible.
Flakiness needs evidence. We keep a simple record of test outcomes on main. A test can be marked flaky only if that record shows it failing and passing on the same commit. The UK postcode test had passed on every run for eight months before that merge. Two minutes with that record would have ruled out flakiness.
Using an agent to investigate is fine. We did not ban agents from red builds. They are genuinely good at reading a failing test’s output, finding the commit that changed the relevant code, and explaining the connection. The owner can use one to diagnose, and often does. What changed is that the agent’s output is a diagnosis and a proposed fix for a human to review, not a commit that goes straight to main. And since that Wednesday, nothing goes straight to main at all: every fix goes through a pull request and its checks, including fixes for red builds.
Closing the gap that caused it
The rules above are about response. We also reduced how often main goes red in the first place.
The biggest change was moving the checkout-related end-to-end tests into the pull request checks. Not the whole forty-minute suite, but the ten minutes that cover checkout, payment and account creation, which are the flows where a failure costs money. The rest still runs after merge. Agent-written pull requests now fail before merge when they break checkout, which is where we want the failure to happen: in front of the person who ran the agent, while they still remember the change.
We also asked engineers to tell their agents about the slower suite. The repository instructions now describe which end-to-end tests cover which areas and how to run them locally, and ask the agent to run the relevant ones before calling a change to shared validation, checkout or authentication done. It does not catch everything, but agents follow that kind of instruction reliably when it is written down.
Finally, we started tracking red-main time as a team metric, reviewed monthly alongside deployment frequency. In the three months since the rules went in, post-merge failures dropped from eleven a month to four, and the median time to green fell from two hours to eighteen minutes, almost all of it through reverts. That is roughly sixty engineer-hours a month back, which is the number I took to my own manager.
The author field is not enough
An agent can write the change, run the tests it knows about, and even diagnose a failure afterwards. What it cannot do is own the consequences on a Wednesday evening when its session has ended and the person who started it has gone home. Somebody has to, and until we said who, the answer was effectively nobody, followed by whoever was most willing to make the red go away.
The rule we ended up with is short enough to fit on a sticky note: whoever merged owns the red build, anyone may revert after fifteen minutes, and no test gets skipped to turn the build green. Most of the effort went into making reverts feel normal. The UK postcode test runs on every pull request now, and it has not been skipped since.