The Agent That Rewrote a File You Did Not Mention

Tariq Oduya

Tariq Oduya

September 30, 2026

The Agent That Rewrote a File You Did Not Mention

The maintainer’s reply was polite and short. “Thanks for the fix to the flag parsing. Could you split this? I can’t review the changes to errors.py as part of a bug fix, and I’m not sure we want them at all.” Then the pull request was marked as waiting on author, where it sat for a week while I felt embarrassed.

I had not asked anyone to touch errors.py. The bug was in how a small open-source command-line tool handled repeated --exclude flags: the second one silently replaced the first instead of adding to the list. I found the cause in the argument parser, described it to a coding agent, and asked it to fix the bug and add a test. It did both, correctly. It also rewrote errors.py: converted every %-style string to f-strings, reordered the imports, added type hints to every function, and, most seriously, changed the base class of two custom exceptions so they inherited from ValueError instead of a project-specific ToolError.

That last change would have broken anyone downstream who caught ToolError. The project has a few hundred dependants. I had glanced at the diff, seen that the parser fix and test were right, noticed a lot of green in another file that looked like tidying, and submitted. The maintainer, who had spent years building trust with those dependants, saw a bug fix with a surprise breaking change riding along.

I have contributed to open source for about a decade. I know better than to bundle unrelated changes. What surprised me was how easy it was to do by accident when a tool did the bundling for me.

Why agents touch files you did not mention

After that week I paid attention to every unrequested file change in my agent sessions, across my own projects and contributions. The causes fell into a handful of patterns.

The file was on the path. To fix the parser, the agent had to understand how errors were raised when a flag was invalid. It opened errors.py to see the exception classes. Once a file is open and in context, it becomes a candidate for editing. An agent that has read a file and noticed things it considers improvable will sometimes improve them, especially if nothing tells it not to.

A type checker or linter complained. In some sessions, the agent ran the project’s type checker after its fix, saw pre-existing warnings in neighbouring files, and fixed those too. From its point of view, it was leaving the codebase cleaner than it found it. From a reviewer’s point of view, it was expanding the change.

Consistency pressure. The fix in the parser used an f-string. errors.py used % formatting. Some models seem to feel a pull towards making nearby code match whatever style they just wrote. The exception hierarchy change came from the same instinct, I think: the agent’s new code raised a ValueError-like exception, and it “harmonised” the existing classes to match.

Vague quality instructions. My global rules file at the time included a line like “write clean, modern, well-typed Python.” That line was meant to govern the code the agent wrote. The agent reasonably read it as applying to any code it touched, and it touched errors.py.

Formatters on save. In one session, the agent’s edits triggered the editor’s format-on-save with a different configuration than the project used, reformatting whole files. Not strictly the model’s doing, but it showed up in the same diff and did the same damage.

A room under renovation where the neighbouring wall has also been painted in a different shade, with a paint roller and tray on a drop cloth

Why it matters more than it looks

It is tempting to see unrequested tidying as harmless. Some of it is. Converting string formatting does not usually break anything. But unrequested changes have costs that are easy to underestimate, especially on a shared codebase.

They hide the real change. A reviewer looking at a diff with forty lines of parser fix and two hundred lines of formatting has to find the forty lines that matter. The more noise, the more likely something important slips past. In my case, the important thing was the exception base class change, buried in a wall of f-string conversions.

They can carry behaviour changes disguised as style. Changing an exception’s base class looks like tidying and is an API break. Reordering imports can change behaviour in modules with import-time side effects. Adding type hints can, in some frameworks, change runtime behaviour. “Cosmetic” is a claim, not a guarantee.

They pollute history. git blame on errors.py would have pointed every line at my “fix repeated –exclude flags” commit. Anyone trying to understand why the exceptions changed would have found a bug fix that says nothing about exceptions.

They create merge conflicts for other people. Two other contributors had open pull requests that touched errors.py. My tidying would have conflicted with both.

They cost trust. In open source especially, maintainers decide whether to engage with a contributor partly on whether their pull requests are easy to review. A fix that arrives with surprise changes teaches the maintainer to look at your future pull requests with more suspicion. That is a real cost, and it lands on the human, not the tool.

What I do before submitting now

The fix that has helped most is embarrassingly simple: I look at the list of changed files before I look at any code.

git diff --stat takes a second. For the parser bug, it would have shown three files: the parser, the test file, and errors.py with a large change count. The first two I expected. The third should have stopped me immediately. I now treat any file in that list that I did not expect as a question to answer before going further: why did this change, and does it belong in this commit?

Usually the answer is that it does not belong. Then I use git checkout -- errors.py to throw the file’s changes away entirely, or git add -p to keep only the hunks that are genuinely part of the fix. It is rare that an unrequested file change is both correct and in scope.

Scissors cutting a long strip of paper into smaller neat pieces on a cutting mat, top-down view

Stopping it before it happens

Catching unrequested changes at review is necessary. Preventing them is better, and a few changes to how I set up sessions have cut them down to a rare occurrence.

Name the scope in the task. Instead of “fix the repeated –exclude bug and add a test,” I now write “fix the repeated –exclude bug in cli/parser.py and add a test in tests/test_parser.py. Do not modify other files. If another file needs to change, tell me why before editing it.” That last sentence matters. Sometimes another file genuinely does need to change, and I want the agent to ask rather than either silently change it or silently fail.

Replace vague quality rules with scoped ones. My rules file no longer says “write clean, modern Python.” It says: “Match the existing style of the file you are editing. Do not refactor, reformat or modernise code outside the lines required for the task. Do not fix pre-existing lint or type warnings unless asked.” Explicitly naming the behaviour I do not want works better than hoping a general instruction will be read narrowly.

Use the project’s formatter configuration, or none. I disabled format-on-save for agent edits in projects where my editor settings differ from the project’s. When a project ships a formatter config, the agent can run it on changed files only.

Keep tidying as its own task. The f-string conversion was not a bad idea in itself. If I want it, I can open a separate pull request, “modernise string formatting in errors.py,” and let the maintainer decide on its own merits. The exception base class change, if anyone wants it, is a breaking change that needs a discussion and a major version, not a ride-along.

When the extra file is genuinely needed

Not every unexpected file in the list is scope creep. Sometimes a fix really does require touching something you did not name. If the parser bug had been caused by a shared helper in utils.py, changing that helper would be the fix, and leaving it alone to keep the diff small would be the wrong call.

The test I use is whether I can explain the extra change in one sentence that starts with “the fix requires.” If I can, the change stays, and I mention it in the pull request description so the reviewer is not surprised. If the sentence starts with “while it was there” or “for consistency,” the change goes into its own branch or into the bin. That one-sentence test has held up across every project I contribute to, and it keeps me honest about whose idea each change was.

What happened to the pull request

I reverted errors.py, force-pushed a branch with only the parser fix and its test, and apologised in the thread for the noise. The maintainer merged it the same day with a thumbs-up. A few weeks later, I opened a separate pull request proposing f-strings in errors.py, without touching the exception hierarchy. It was merged too, after a short discussion about minimum Python version.

The exception change never went in. Nobody wanted it. The agent had made a design decision about a public API because it seemed more consistent, and I had almost shipped it under a bug-fix title without noticing.

I do not think the agent was wrong to read errors.py. It needed to understand the exceptions to fix the parser properly. The problem is that reading and editing feel like the same activity to the tool, and they are very different activities to everyone who has to live with the result. Keeping them separate is now part of my job whenever I use one, and the file list in git diff --stat is where I do it.

More articles for you