Rate Limits and Quota Errors: What the Agent Does When the Model Stops Mid-File
Quinn Reed
September 30, 2026
I gave the agent a job I thought was safe to leave: replace our homegrown logger with structured logging across eleven service files. Same change in every file, clear pattern, tests in place. I wrote the task, watched the first two files go through, and went to lunch.
When I came back, the terminal showed a rate-limit message, a few retries with growing waits, and then a flat “usage limit reached” that ended the session. Six files had the new logger. Five had the old one. One of the six, the payments service, stopped in the middle of a function. The last line in the file was half a logger.info( call, and the closing braces for the class were gone.
None of that was a disaster. It was a Tuesday afternoon of cleanup. But it taught me that “the model stopped” covers three different events, each leaves the repository in a different state, and the agent has no idea which one happened when it starts again.
Three ways a model stops
A rate limit. Too many requests or tokens per minute. The provider returns an error, often an HTTP 429, or an “overloaded” response when it is busy on its end. This is temporary. Most agent tools retry automatically with a growing delay, and if the retry succeeds, the session carries on as if nothing happened. The danger is small, as long as the retry resends the same request rather than a new one.
A quota or spend cap. You have used your allowance for the day, the month, or the budget someone set on the account. Retrying does nothing until the limit resets or someone raises it. The session ends. Whatever plan the agent was holding in its context is gone, and the next session starts with nothing but the files on disk.
An output length cap. This one is not an error at all, which is why it is the nastiest. Every model has a maximum number of tokens it will produce in one response. If the agent asks it to write out a whole large file and the file is longer than that cap, the response simply ends where the budget ran out. The API reports it politely, with a finish reason along the lines of “length” instead of “stop”, and a tool that does not check that field will write the truncated text to disk as if it were finished.
My half-finished payments file was the third kind hiding inside the second. The agent had been rewriting that file in full rather than editing it, the response was cut off at the output limit, the tool wrote what it had, and then the quota ran out before the agent could run the tests and notice. Two unrelated limits, one after the other, and the repository showed only the combined damage.

What the repository looks like after each
It helps to know what state to expect, because the recovery is different.
After a rate limit that recovered, nothing is wrong. The agent picked up where it was. If you were not watching, the only sign is that the task took longer.
After a quota stop, you have a partial set of complete edits. Most agents apply changes one tool call at a time, so each file the agent touched is either fully changed or untouched. The code in any single file compiles. The codebase as a whole may not, because half the call sites use the new interface and half use the old one. In my case, six services imported a logger module that five others did not know about, and a shared helper had been updated for the new signature, which broke the five that had not been converted.
After an output cap hit on a whole-file write, you have a file that is syntactically broken. The good news is that the type checker or the test runner will scream immediately. The bad news is that if the session also ended, nothing ran them.
After an output cap hit on a normal edit, it depends on the tool. Some reject a malformed edit block and ask again. Some apply what they can. It is worth finding out which yours does before you need to know.
What the agent does when you restart it
This is the part that cost me the afternoon.
I started a fresh session and typed, roughly, “the previous session was interrupted, please finish the logging migration.” The new session did what I would have done with no context: it read the repository. It found six files using the new logger and five using the old one, plus a broken payments file.
Then it made a judgement call. It decided the codebase was “inconsistent” and that the safest fix was to make it consistent. Since the old logger was still used in the shared test utilities and in the majority of the rest of the repository outside those eleven files, it concluded the old logger was the standard and started converting my six migrated files back. It repaired the payments file by reconstructing the old version from the other files’ patterns.
From its point of view, this was reasonable. It had no plan, no record of intent, and a repository that looked like someone had half-done something questionable. My one-line prompt said “finish the migration”, but did not say in which direction, and “migration” is ambiguous when the evidence points both ways.
I stopped it after three files, reset them, and wrote a much longer prompt. That fixed it. But the lesson was clear: the agent does not remember being interrupted. It sees a snapshot, and it guesses the story. If the snapshot is ambiguous, the guess can go either way.
What I changed
None of these are clever. They are what I should have been doing already.
Commit a checkpoint before any long task. I start long jobs on a branch from a clean working tree. That way, “reset to where I was before lunch” is a single command rather than an archaeology project. It also makes the damage visible: git diff --stat shows exactly which files the interrupted session touched.
Break long tasks into steps that each leave the build green. Eleven files in one go was greedy. Now I would ask for the shared logger module and helper first, with the old interface still working, then convert services in small batches, running the tests after each. If the session dies after batch two, the repository is in a working state with a clear boundary between done and not done.
Keep the plan in a file, not only in the chat. For anything longer than a few steps, I ask the agent to write a short checklist to a scratch file in the repository at the start and tick items off as it goes. It costs a few lines. When a session dies, the next one reads the checklist and knows the direction, what is done, and what is next. The file gets deleted before the branch merges. This one change would have saved my afternoon on its own, because “convert the remaining five services to the new logger, see TASK.md” leaves nothing to guess.
Prefer edits over whole-file rewrites for large files. Most agent tools let you steer this. A targeted edit to a 900-line file produces a few dozen lines of output and is nowhere near the output cap. A full rewrite of the same file may be over it. Where the tool insists on full rewrites, I check file length first and split the file if needed. The payments service was overdue for splitting anyway.
Run the checks after every file, not at the end. I added an instruction to our agent rules: after editing a file, run the type checker for that package. A truncated file then fails within seconds of being written, while the session is still alive to fix it, rather than being discovered by whoever opens the repository next.

Handling the limits themselves
The limits are not going away, and some are there to protect you. A few habits make them less disruptive.
Know which limit you have. Rate limits are per minute; quotas are per day or month. They produce different messages. When I hit one now, the first thing I do is read the exact error, because it tells me whether waiting two minutes will help or whether the session is over.
Do not fight the backoff. When the agent is retrying a rate limit, starting a second session to “keep things moving” draws from the same allowance and makes both of them slower. I learned this by doing it. Let the retries run.
Put a warning in front of the hard cap. We keep a monthly spending cap on the team account, and I am not removing it. But we added an alert at 80 percent, so a big task does not start with a few percent of budget left and die halfway through. If you know you are close, start the long job tomorrow, or do the small batches today.
Watch the finish reason if you build your own tooling. If you script against the API directly, check why each response ended. Treat anything other than a normal stop as a failure: do not write the output to disk, and either continue the response or split the request. This is one line of code, and it is the line most homegrown agent scripts leave out.
My recovery routine
When I come back to a stopped session now, I do the same four things before letting any agent near the repository again.
- Read the last error in full and identify which limit it was.
- Run
git statusandgit diff --statto see exactly what changed. - Run the type checker and tests to find broken files.
- Decide the direction myself: finish forward or reset. Then tell the next session that decision in plain words, with the checklist if there is one.
The fourth step is the one that matters. An interrupted agent leaves a half-finished change, and a half-finished change is ambiguous. A fresh session will resolve the ambiguity by guessing, and it may guess the opposite of what you meant. Deciding the direction is a thirty-second human job, and it is the part I no longer hand to the model.
The limits themselves are an inconvenience. The real risk is a new session confidently finishing a job you never asked for, from a snapshot it had to interpret.