When to Stop a Coding Agent and Take the Keyboard Back
Owen MacAllister
September 30, 2026
The seventh edit to SettlementBatchProcessor.java was when I finally noticed what I was doing. I was sitting with my hands in my lap, watching a coding agent try to fix a failing integration test in a fourteen-year-old payments monolith, and I had been watching for fifty minutes. Each attempt came with a fresh, reasonable explanation. Each attempt edited the same three methods. The test kept failing with the same assertion, one cent off on a batch total.
I knew what the bug was around minute twenty. I had seen this codebase round currency in two different places before. I kept waiting because the agent seemed close, and because some part of me had decided that typing the fix myself would be admitting the tool had failed. That is a bad reason to keep a process running. I stopped it, opened the file, found the second rounding call in a utility class the agent had never opened, and fixed it in four minutes.
I have been working with legacy codebases for fifteen years and with coding agents for about two. The skill I had to learn fastest was not prompting. It was noticing the moment an agent session has stopped converging and taking the keyboard back before that moment costs an hour. Here is what I watch for now.
Signal one: the same file, edited again
The clearest sign is repetition. An agent that is making progress tends to move through a problem: it reads, it changes something, the failure changes, it follows the new failure somewhere else. An agent that is stuck circles. It edits the same method with a slightly different approach, runs the same test, gets the same failure, and tries a third variation.
On the settlement bug, the edits were variations on one theme: change the RoundingMode, change the scale, move the rounding before the sum instead of after. All of them were plausible fixes for a rounding bug in that method. None of them could work, because the extra rounding happened elsewhere, in MoneyUtils.normalise(), called from a mapper two layers down.
My rule now is simple. If an agent edits the same region of code three times for the same failure and the failure does not change, I stop it. Not because the fourth try cannot succeed, but because three failures on one hypothesis mean the hypothesis is probably wrong, and the agent is not going to abandon it without a nudge.
Signal two: the failure stops changing, but the explanation keeps growing
This one took me longer to notice. When an agent is stuck, its explanations often get longer and more elaborate while the evidence stays the same. The first attempt says “the total is off because of rounding mode.” The fifth says something about floating-point accumulation interacting with the batch partitioning strategy and the order of settlement records under concurrent processing.
The code does not use floating point. There is no concurrency in that test. The explanation had become a story built to fit a failure it could not explain, and it sounded more expert as it got less true. A growing explanation over a static failure is not a sign of deeper understanding. It usually means the model is reasoning from its own previous guesses rather than from what the code does.

Signal three: the scope starts widening on its own
Sometimes a stuck agent does not circle. It expands. It decides the problem is architectural and starts restructuring: extracting a class, changing an interface, adding a new abstraction to “centralise” the logic. In a greenfield project that might occasionally be the right call. In a legacy codebase, where one interface has forty-seven implementations and some of them are loaded by reflection from a config file, it is how a failing test turns into a broken release.
I watch the file list. If I asked for a bug fix in one service and the diff now touches shared utilities, interfaces, or build configuration, I stop and ask why. Often there is a good reason. Just as often the agent has lost the thread and is trying to make the problem go away by changing the ground it stands on.
Signal four: it starts negotiating with the test
The most dangerous signal is when the agent turns its attention from the code to the test. It updates the expected value. It adds a tolerance to the assertion. It marks the test as flaky and adds a retry annotation. It mocks out the component that produces the wrong answer.
Each of these can be legitimate. A test with a wrong expected value exists. Floating-point comparisons do need tolerances. But when the test has passed for years and started failing after a change, the test is almost never the problem. I stop any session the moment it proposes changing an assertion I did not ask it to question. If you let this one run, you do not get a stuck session. You get a green build that ships a bug, which is much worse.
Signal five: you already know the answer
This is the signal I ignored for thirty minutes on the settlement bug, and it is the most human one. If you can see the fix, stop and type it. Waiting for an agent to rediscover something you know is not delegation. It is spectating.
There is a pull, especially early on with these tools, to let the agent find it. Partly curiosity about whether it can. Partly a feeling that the point of the tool is to not do the work yourself. Both are reasonable in a sandbox and expensive on a Thursday afternoon with a release branch cut at five. Your knowledge of the codebase is the most valuable context in the room. If the agent cannot get to it, your hands can.
Signal six: the session has gone long
Long sessions degrade. After enough back-and-forth, the context is full of failed attempts, long test outputs, and the agent’s own previous explanations. In my experience the quality of reasoning drops noticeably after a session has accumulated a large pile of failed attempts. It starts repeating earlier ideas as if they were new, or forgets a constraint you gave it at the start.
If a session has been running for more than about half an hour on a single problem without clear progress, I stop it even if none of the other signals has fired. Sometimes I restart with a fresh session and a much tighter description built from what I learned. Sometimes I take over. What I no longer do is add one more message to a session that is already carrying an hour of wrong turns.

How to take the keyboard back without losing the good parts
Stopping an agent is not throwing away its work. Even a stuck session usually produced something useful: it narrowed the search, it wrote a reproduction, it ruled out a few hypotheses. The way you stop matters.
Stop, then look at the whole diff. Before touching anything, I run git diff and read everything the agent changed. On the settlement bug there were edits to three methods, a new test helper, and a changed log line. The test helper was genuinely useful. The rounding edits were all wrong.
Keep what is right, revert what is not. I use git add -p to stage the parts worth keeping and discard the rest. Starting your fix on top of three failed attempts is how you end up with code nobody can explain six months later. In a legacy codebase, that is how the next rounding bug gets born.
Write down what the agent ruled out. The failed attempts were not worthless. They told me the problem was not in the rounding mode or scale inside the processor. That pushed me outward toward callers and utilities. I note these in the ticket, because the next person debugging this code, possibly me, will benefit from knowing which theories were already tested.
Fix the thing yourself, then hand back the tedious parts. After I found the double rounding in MoneyUtils, there were eleven other call sites that needed the same change and six tests that should cover them. That part went back to the agent, with a specific description: remove the normalise() call at these call sites, add a test at each one asserting the unrounded intermediate value. It did that in a few minutes and got it right, because the thinking was already done.
Why the agent got stuck in the first place
It is worth asking why a session failed, because the answer usually tells you how to set up the next one.
On the settlement bug, the agent never opened MoneyUtils. It was two calls away from the failing method, through a mapper with a generic name, and nothing in the test or the processor mentioned it. The agent’s search stayed where the symptom was. A human who had not worked in this codebase would have made the same mistake. I made it myself, eight years ago, the first time I met this code.
Afterwards I added a paragraph to our agent instructions file: currency rounding in this system happens in MoneyUtils.normalise() and in the settlement processor; always check both. That kind of line is worth more than any amount of prompt cleverness. The session did not fail because the tool was bad. It failed because the knowledge that mattered lived only in my head, and the agent could not get into my head.
Stopping is part of using the tool well
There is an idea floating around that the ideal agent workflow is hands-off: describe the task, walk away, come back to a finished change. For some tasks, that works. In a codebase where half the business logic is folklore and the other half is in stored procedures, I spend a fair amount of my agent time with my hand near the stop button.
That is not a failure of the workflow. It is the workflow. The agent is very good at reading many files quickly, trying hypotheses, and doing repetitive changes without getting bored. I am good at knowing where the bodies are buried. The sessions that go well are the ones where each of us does our part, and the handoff happens as soon as the agent’s part stops producing progress.
I still think about those fifty minutes. The fix took four. The expensive part was not the bug, and not the tool. It was my reluctance to notice that I had stopped collaborating and started watching. These days, if I catch myself sitting back with my arms folded while an agent edits the same method for the third time, I take that as the signal. Stop it, read the diff, keep what is good, and type the fix.