Two Agents in One Repo: The Merge Conflict Neither of Them Can See
Owen MacAllister
September 30, 2026
Both pull requests were green. Both had been reviewed. They merged twenty minutes apart on a Thursday afternoon without a single conflict. By Monday morning, a customer who had been promised a hundred and fifty dollars of credit for the previous week’s outage had received a dollar fifty, and so had everyone else who had been given a goodwill credit since Friday.
Nobody had written a bug, exactly. Two coding agents, working in parallel on two branches for two engineers on my team, had each made a correct change. The problem existed only in the combination, and nothing in our process ever looked at the combination before it reached production.
I have worked on large codebases for fifteen years and have seen this kind of failure before, between people. What was new was how easily we walked into it once each engineer could run two or three agent sessions at once.
What each agent did
The first agent was migrating the billing module from storing amounts as decimal dollars to storing them as integer cents. It is a common, sensible change, and this one was careful. It updated the database columns with a migration, changed every function in the billing module to take and return cents, updated all existing call sites, and updated the tests. Among the functions it changed was applyCredit(accountId, amount), which now expected amount in cents.
The second agent, started the same morning from the same commit on main, was building a new feature: support staff could grant a goodwill credit from the admin panel. It added a form, an endpoint, an audit log entry, and a call to applyCredit(accountId, amount) with the amount the support person typed, in dollars. It wrote tests that granted a credit of 1.50 and checked the balance went up by 1.50. On its branch, where applyCredit still took dollars, all of that was correct.
The two branches touched different files. The billing migration never touched the admin panel, because the admin panel’s credit feature did not exist when that agent started. The credit feature never touched the billing module’s internals; it only called a public function. Git saw no overlapping lines and merged both without complaint.
After the merge, the function signature was identical, the argument was still a number, and the type checker had nothing to say. A support person typed 150 for a hundred-and-fifty-dollar outage credit, the endpoint passed 150 to applyCredit, and applyCredit added 150 cents. Credits with a decimal part, like 12.50, failed a validation check in the new billing code and showed the support person a generic error, which they assumed was a glitch and retried with a round number. Every credit granted over those four days was a hundredth of what the customer had been told. Nobody noticed until a customer wrote in to ask why their outage apology was worth less than a coffee.

Why neither agent could see it
It is tempting to say the agents should have noticed. I do not think they could have, in the setup we gave them.
Each agent works from a snapshot. It sees its own branch: the commit it started from, plus its own changes. It does not see the branch another agent is working on in another session, and it has no idea that another session exists. From inside either branch, the code was consistent and the tests told the truth.
Git’s conflict detection is textual. It notices when two changes edit the same lines. It has no concept of “this function’s meaning changed” or “this new code relies on the old meaning”. A semantic conflict, where two changes are each valid but incompatible in meaning, merges cleanly every time.
People working in parallel have the same blind spot, but they have some defences an agent lacks. They talk at stand-up. They see each other’s pull requests in the team channel. An engineer migrating money to cents would probably mention it, and an engineer adding a credit feature would probably hear. Two agents started by two people on the same morning have none of that, and the people who started them were each reviewing their own agent’s diff, not thinking about the other one.
The same-directory version is worse
Before this incident, we had already learned a cruder version of the lesson. Early on, one engineer ran two agent sessions in the same checkout, one fixing a bug and one writing tests elsewhere. Within an hour, one session had run the formatter across the project and rewritten files the other was halfway through editing. The other session then ran a git checkout of a file to undo its own mistake and discarded the first session’s work on that file.
Neither agent knew the working tree was shared. Each assumed that any change it had not made itself was either pre-existing or its own earlier mistake. That rule is now absolute on our team: every parallel agent session gets its own working copy, either a separate clone or a Git worktree. Worktrees are cheap and make this easy.
But separate working copies only solve the collision problem. They make the semantic conflict problem slightly worse, because each agent is now even more cleanly isolated from what the others are doing.

What we changed
Test the combination, not just the branches. The single most effective change was also the most boring. Our branch protection now requires a pull request to be up to date with main before it can merge, so CI runs against main plus the change, not against the stale main the branch started from. On a busy repository, that turns into a merge queue: pull requests are tested in the order they will land, each on top of the ones ahead of it. With that in place, the credit feature’s tests would have run against the cents-based applyCredit and failed immediately. The test that granted 1.50 would have been rejected by the cents validation before anyone merged it.
Make semantic changes break loudly. The deeper problem was that a change in meaning left the signature untouched. We now follow a rule for any change in the meaning of a function’s inputs or outputs: change something the compiler can see. The billing code now uses a Cents type that plain numbers cannot be passed to without an explicit conversion. Had that existed, the credit feature would have failed to compile after the merge, and the type checker in CI would have caught it even without the up-to-date rule. Where a type change is impractical, renaming the function, say to applyCreditCents, does the same job: any code written against the old name stops compiling. Deliberately turning a semantic conflict into a visible one is one of the oldest tricks in large-codebase work, and it matters more when changes are written in parallel by processes that cannot talk to each other.
Contract changes go first, and alone. We now treat changes to shared contracts, such as public functions of a module, database schemas, API shapes and event formats, as a separate kind of work. They are done in one session, merged, and only then do other sessions start work that depends on that module. Running a contract change in parallel with feature work that uses the contract is the exact setup that produced our bug. Fanning out after the contract is settled is fine.
Ask the agent to check what moved on main. Before any agent-written pull request is marked ready, the agent rebases it onto the latest main and answers one question in the description: which functions, types or schemas that this change depends on have been modified on main since the branch started? Agents are good at this when asked. They can read the log and the diffs of the files they call into, and they will say things like “applyCredit now expects cents; I’ve updated my call to convert”. It is not a guarantee, but it moves the check to the one participant who has time to do it thoroughly.
Keep one list of what is in flight. We keep a short board of active agent tasks, each with one line on which modules and contracts it touches. It is maintained by people, not agents, and it takes seconds to update. Its only job is to make someone notice when two tasks touch the same contract, so they can be ordered instead of run side by side.
When running agents in parallel is fine
None of this means parallel agents are a bad idea. They are one of the genuinely useful things this tooling makes possible, and we still run several at once most days.
Work parallelises well when the pieces do not share a contract: separate leaf features in different parts of the product, test coverage for existing code, documentation, dependency updates in isolated packages, and exploratory spikes that will not be merged. Work parallelises badly when one task changes the meaning of something that another task uses, and it is especially bad when that meaning change leaves the types and names untouched.
The question I ask before starting a second session is simple: if both of these land, is there any function, table or payload whose meaning one of them changes and the other relies on? If the answer is yes, they run one after the other. If I am not sure, they run one after the other.
The conflict is between assumptions
We corrected the credits within a morning: a query to find every goodwill credit since Friday, a recalculation, and an apology email that support wrote better than I would have. The code fix was one line. The process fixes took a week to settle.
What stays with me is how normal both pull requests looked. Each was a clean, well-tested change that any reviewer would approve. The bug lived in the gap between two sets of assumptions, one agent’s about what applyCredit meant after its migration and the other’s about what it meant before. Git compares lines, not assumptions, and so do most review processes. When the authors can no longer overhear each other, something in the pipeline has to test the combination and make changes of meaning visible.