Context Window vs a Fresh Chat: When Stuffing the Repo Costs More Than It Helps
Elena Vasquez
September 30, 2026
The rate limiter bug should have been a fifteen-minute fix. Our API gateway was letting a small number of requests through after a client hit its limit, roughly one extra request per window under load. I knew it was a race in the Redis script. I wanted the agent to find it and propose a fix, and because I wanted it to “really understand the system,” I attached the whole services/ directory to the conversation. Eleven services, a few hundred files. Modern context windows are huge. Why not use them?
The agent came back with a fix for a rate limiter in a different service. Our billing service has its own, older limiter using a token-bucket implementation with a similar name, and the agent had blended the two: it diagnosed the gateway’s sliding-window bug using the billing limiter’s config structure, and proposed a patch that referenced a setting the gateway does not have. It was fluent, confident, and stitched together from two places.
I started a fresh conversation, attached three files (the gateway’s limiter module, its Lua script, and the failing test), and described the symptom in two sentences. The agent found the race, a GET followed by a separate INCR instead of doing both atomically in the script, on its first answer.
That afternoon broke a belief I had held since context windows passed a hundred thousand tokens: that more context is always at worst neutral. It is not. Past a certain point, more context makes the answers worse, and knowing when to stop adding and start fresh has become one of the more useful habits I have.
Why more context can make answers worse
I am a backend engineer, not a machine learning researcher, so I will describe this in terms of what I observe rather than what happens inside the model. But the observations are consistent enough that I plan around them.
Similar things blur together. The billing limiter and the gateway limiter had similar names, similar concepts and overlapping vocabulary. With both in context, the agent treated them as one system. Any large codebase has these near-duplicates: two HTTP clients, three date helpers, a v1 and a v2 of the same handler. Attach everything and you invite the model to average them.
The relevant part gets less attention. The gateway’s Lua script was about 40 lines in a context of several hundred thousand tokens. It was there, technically. But being present in the context is not the same as being weighed properly. Models are generally better at using information that is prominent and close to the question than information buried in the middle of a very long input.
Irrelevant code suggests irrelevant fixes. Show a model eleven services and ask for a fix, and it will look for patterns across all eleven. Sometimes that is useful. More often it produces a “consistent” fix that borrows conventions from a service that has nothing to do with the problem.
None of this means large context windows are useless. They are excellent when the task genuinely spans a lot of code: a cross-service rename, an audit of how every service handles a header, a summary of a module you have never read. The mistake is treating the window as a place to dump everything that might conceivably matter, instead of a place to put what does.

The long session is the same problem, slower
Attaching a whole directory is the fast way to overfill a context. The slow way is a long session.
A week after the rate limiter episode, I spent most of a morning in one conversation refactoring our request validation layer. Early on, the agent proposed moving validation into a middleware; I said no, we keep it per-handler for now, because some handlers need raw bodies. We moved on. Seventy minutes and dozens of edits later, while fixing an unrelated test, the agent reintroduced a shared validation middleware as part of its fix.
It had not forgotten my instruction in any simple sense. The instruction was still there, far back in the conversation. But it was now surrounded by a morning of file contents, test outputs, abandoned attempts and its own earlier proposal of the middleware, which was also still in the context, written out in detail. The rejected idea and the rejection were both present. The idea was longer and more concrete.
This is the part of long sessions that I underestimated. A conversation accumulates not just useful history but also dead ends, superseded code, and discarded plans. Each of those is context that pulls on future answers. The session becomes a record of everything you tried, not just what you decided.
Signs a context has gone bad
These are the signals that make me stop adding to a conversation and start fresh:
- It mixes up similar things. Config keys from the wrong service, a helper from the wrong module, a function signature from the old version of a file it edited earlier.
- It revives rejected ideas. If something I explicitly turned down comes back, the context has enough weight on the wrong side.
- It refers to code that no longer exists. After several edits, it quotes a version of a function from an hour ago instead of the current one.
- Answers get vaguer. Early answers cite specific lines. Later ones talk about “the validation logic” and “the relevant handler” without naming them.
- I find myself repeating constraints. If I have restated the same rule three times in one conversation, the conversation is the problem.
What a good fresh start looks like
Starting a new conversation is not the same as starting over. The work done in the old session is in the files. What needs to carry over is the small amount of decision-making that is not in the code yet.
When I end a long session now, I ask the agent to write a short handoff: what the task is, what has been done, what was decided and why, what is still left. Then I read it and edit it, because the agent’s summary tends to include too much and sometimes gets a decision wrong. The version I carry forward is usually five to ten lines.
For the validation refactor, it looked like this:
- Task: move request validation from ad-hoc checks to schema objects, handler by handler.
- Done: users, orders, and webhooks handlers converted and tested.
- Decided: validation stays per-handler, not middleware, because the webhooks and uploads handlers need raw bodies.
- Left: payments and uploads handlers.
- Files to start from:
handlers/payments.py,handlers/uploads.py,validation/schemas.py.
That fresh session finished the remaining two handlers in about twenty minutes with no middleware in sight. The decision was stated once, near the top, and nothing in the context argued with it.

Letting the agent choose what to read
The other habit I dropped was pre-loading context at all when I did not have to. Most coding agents can search the repository themselves. If I describe the symptom well and point at an entry file, they will usually find the related code by following imports and searching for names.
That has a cost, since exploration takes time and turns, but it has a real advantage: the agent pulls in what the problem actually touches, one piece at a time, rather than everything I thought might be relevant. When I attached the whole services directory, I was making the relevance decision on the agent’s behalf, badly. When I give it three files and a description, it opens a fourth or fifth only if it needs them.
This works less well in repositories with confusing names or duplicated modules, where the agent’s search lands on the wrong candidate, and the reasons come down to how coding agents actually find your code: names, paths, and whatever the index can see. In those cases I still attach specific files, just not whole directories. The goal is a small, precise starting point, not a complete one.
There is also a bill attached to stuffing. Everything in the context is re-sent every time the agent takes an action, which is most of what a long agent task quietly spends, so a session that starts with several hundred thousand tokens of attached code is expensive before it has done anything useful. I care about that, but in my experience the quality problem shows up first. By the time the cost is noticeable, the answers have usually already started to drift.
My working rules
Start small, grow on demand. Attach the files I am confident matter. Let the agent pull in the rest. If I think “this might be relevant too,” that is usually a sign it is not.
Keep near-duplicates apart. If there are two similar implementations and only one is involved, I make sure only that one is in the context. If both are involved, I say so explicitly and name which is which.
One conversation, one task. When a task is done, I start a new conversation for the next one, even if it is related. A bug fix and the refactor it inspired are two sessions.
End long sessions deliberately. Around the hour mark, or at the first sign of confusion, I ask for a handoff, edit it, and start fresh.
Use big context for genuinely big questions. “How does authentication flow across these four services?” is a good use of a large window. “Fix this bug” rarely is, no matter how important the bug.
The window is not the point
Context windows will keep growing, and each new model is better at using long inputs than the last. I expect some of the specific failures I have described to get less frequent. But I do not expect the underlying trade to go away. Relevant context helps. Irrelevant context competes. A long history of dead ends is not neutral.
The rate limiter fix, for what it is worth, was a four-line change: move the GET and INCR into a single Lua call. The agent found it in one pass once it could see only the code that mattered. I spent more time attaching the wrong eleven services than it spent finding the right forty lines.