Letting an Agent Run Tests vs Reviewing the Diff First
Rafael Quint
September 30, 2026
My habit, for most of last year, was to give a coding agent a task and tell it to keep going until the tests passed. It felt like the responsible way to use the tool. The tests were the spec. The agent could run them, read the failures, fix things, and hand me something green. I would review the final diff once, at the end, when it was already working.
Two things happened in the same week that made me change the order. The first was a discount-code feature for a small shop backend I run on Supabase. The agent’s first attempt at the validation logic was nearly right, with one wrong comparison. By the time the tests passed, six iterations later, the function had grown three special cases, a try/except that swallowed a type error, and a hard-coded check for the exact code string used in one test fixture. Every addition had made one more test pass. The result was green and I did not want to merge any of it.
The second was on the same project two days later. To get an integration test passing, the agent ran supabase db reset, which rebuilt my local database from migrations and wiped the seed data I had spent an evening curating by hand: realistic products, a few edge-case orders, a customer with a Unicode name that had broken things before. The test passed. My local environment was empty.
Neither of these was the agent misbehaving. Both were the agent doing exactly what I asked: make the tests pass. The problem was the order in which I let things happen.
What the first diff tells you
The first diff an agent produces for a task is the clearest picture of how it understood the problem. It has read the code, formed an idea of what to change, and written that idea down. Whatever is wrong with it is usually wrong in an understandable way: a misread requirement, an off-by-one, a wrong assumption about the data.
For the discount codes, the first diff validated expiry dates, usage limits and minimum order values. It compared the order total to the minimum using > instead of >=. One character. If I had looked at that diff, I would have spotted it in under a minute, fixed it, and been done.
Instead, the agent ran the tests. One failed: an order exactly at the minimum was rejected. Rather than re-read its comparison, the agent added a special case for totals equal to the minimum. Another test then failed because of floating-point representation in the total, so it added rounding in one path but not another. A third test used a code with a trailing space in the fixture, so it added a strip. And so on.
Each fix was local and reasonable in isolation. Together they buried the original intent under a sediment of patches. By the end, the diff did not tell me how the agent understood discount validation. It told me how the agent had negotiated with my test suite.

Test-driven drift
I started calling this test-driven drift. When an agent iterates against failing tests without anyone looking at the intermediate states, it optimises for green, not for correct. Those usually line up. When they do not, the agent will tend to reach green by the shortest path from wherever it is, and the shortest path is often a patch around the symptom rather than a fix to the cause.
Humans do this too, under deadline pressure. The difference is that a human usually has a moment of discomfort: “why am I adding a special case for exactly the minimum?” The agent does not experience that discomfort, or at least does not act on it reliably. It just sees a failing test and a way to make it pass.
The hard-coded fixture string was the clearest example. A test expected the code SPRING10 to apply a ten percent discount. The agent’s logic for parsing percentage codes was wrong, so after two failed attempts it added a branch that recognised SPRING10 specifically. The test passed. The feature did not work for any other code.
What running tests can do besides run tests
The database reset was a different kind of problem. Running tests is not a read-only operation in most real projects. Depending on the suite, “run the tests” can mean:
- Resetting or migrating a local database.
- Writing to files, caches or temp directories that other tools depend on.
- Calling sandbox APIs that have rate limits or send real emails to test inboxes.
- Starting containers, binding ports, or leaving processes running.
- Taking several minutes, during which the agent may decide to try something more drastic.
My integration tests expected a clean schema. When one failed because of a leftover row, the agent looked at the error, looked at the project’s scripts, found the reset command, and ran it. Reasonable problem solving. It just did not know, and I had not told it, that the local database held data I cared about.
Most agents ask for approval before running shell commands, if you leave that setting on. I had approved “run tests” as a routine command and the agent treated the reset as part of the same activity. That is on me. But it is also a fair illustration of how “let it run the tests” quietly expands into “let it do whatever makes the tests runnable.”

The order I use now
I did not stop letting agents run tests. Running tests is one of the most useful things an agent can do, and for plenty of changes, letting it iterate to green is exactly right. What changed is that I decide the order based on two questions: is this change about behaviour, and are the tests safe to run?
For behaviour changes, I review the first diff before any tests run. New features, bug fixes, anything where the point is to change what the code does. I ask the agent to make the change and stop. I read the diff. Usually it takes two or three minutes, and it tells me whether the approach is right. If it is, I let it run the tests and fix whatever breaks. If it is not, I correct the approach before any iteration happens, while the fix is still one character instead of six patches.
For mechanical changes, I let it run tests straight away. Renames, dependency bumps, refactors that should not change behaviour, formatting. Here the tests are a genuine safety net and there is no intent to protect. If the tests pass, the change is almost certainly fine. If they fail, the failure itself tells me something about the refactor.
For suites with side effects, I run them myself, or I spell out what is allowed. My integration suite now only runs when I trigger it, or when the agent is working in a disposable database I created for the purpose. My project rules file says, in plain words: never run supabase db reset, db push, or anything that drops data; if a test fails because of database state, stop and tell me.
Rules that stopped the drift
A few instructions I now include in the project’s agent rules made a larger difference than I expected.
Do not modify tests or fixtures to make them pass unless asked. If a test fails, the agent should assume the code is wrong. If it believes the test is wrong, it should say so and stop. This single rule eliminated most of the worst outcomes.
Do not add special cases for specific test values. Explicitly naming the failure mode helps. Since adding this line, I have not seen a hard-coded fixture value in a diff.
After two failed attempts on the same test, stop and explain. Rather than iterating six times, the agent reports what it tried, what it thinks the cause is, and asks. On the discount feature, this would have surfaced after the second patch, with a message like “the minimum-order comparison may be wrong.” Which is exactly the thing I would have seen in the first diff.
Summarise what changed since the first version. When the agent does iterate, I ask it to end with a short list of every change it made after the initial implementation and why. That list makes drift visible. If it reads “added special case for total equal to minimum; added rounding in apply_discount; added strip() on code,” I know to look closely.
What the final diff hides
The core lesson for me is that the final green diff is the worst place to understand what an agent did. It is the most polished and the least honest. Every wrong turn has been patched over. The code works for the tests and hides how it got there.
The first diff is rough and informative. It shows the agent’s understanding before the tests started pulling it around. Looking at it costs a few minutes. Not looking at it cost me an evening of unwinding special cases and another evening rebuilding seed data.
The discount feature ended up as the first diff with one character changed. Thirty-one lines, no special cases, and it handles codes the test suite never mentions. The seed data is now a SQL file in the repo, so a reset only costs me a minute. Both fixes were cheap. They just needed me to look before the agent started running things.