Token Use on a Coding Agent: What a Long Task Quietly Spends

Tom Reeves

Tom Reeves

September 30, 2026

Token Use on a Coding Agent: What a Long Task Quietly Spends

I found out what a long agent task costs the way most people do: from a usage page, a few days after the fact. One afternoon I had asked a coding agent to migrate a side project’s test suite from Jest to Vitest. About two hundred test files, some custom matchers, a handful of module mocks that relied on Jest-specific behaviour. It took roughly two and a half hours of wall-clock time, with me dipping in and out. It worked. The pull request was clean.

When I looked at the usage breakdown later that week, that single session had consumed more tokens than the previous ten days of my normal agent use combined. Around fourteen million input tokens and a quarter of a million output tokens. The output, the actual code the agent wrote, was a rounding error. Nearly all of the spend was the agent reading, and mostly reading things it had already read.

I write about developer tool pricing for a living, and I still had the wrong mental model. I thought of token cost as roughly proportional to how much code gets written. For a long agent task, it is closer to proportional to how long the conversation is multiplied by how many times the agent acts. Here is where the tokens actually went, and what I changed.

The loop is the meter

A coding agent works in a loop. It reads the conversation so far, decides on an action (open a file, run a command, edit something), gets the result, and then does it again. Each turn of that loop is a new request to the model, and each request carries the whole conversation: your instructions, any rules files, every file it has opened, every command output it has seen, every edit it has made.

That is the part that surprised me when I first did the arithmetic. If a session’s context sits around 80,000 tokens and the agent takes 150 actions, the model processes something like twelve million input tokens across those requests, even though the “new” information in each step might be a few hundred tokens. The conversation is re-sent every time the agent acts.

Most providers soften this with prompt caching: the unchanged prefix of a conversation is billed at a much lower rate when it is sent again shortly afterwards. That is why a long session does not cost what the raw token count suggests. But caching lowers the price of re-reading; it does not stop it from happening. It also has edges. Caches expire after a few minutes of inactivity, so if I walked away to make tea and came back, the next request paid full price for the whole context again. And anything that changes early in the conversation invalidates the cache for everything after it.

An old analog electricity meter with a spinning dial mounted on a brick wall

Where my fourteen million went

I went back through the session transcript and roughly categorised what filled the context. The proportions are approximate, but the shape was unmistakable.

Test runner output: the biggest single item. Every time the agent ran the test suite, the full output came back into the conversation. Vitest’s default reporter printed every test name, every pass, and full stack traces for failures. A full run of two hundred files produced tens of thousands of tokens of output. The agent ran the whole suite eleven times. Each of those outputs then stayed in the context for every subsequent turn, re-sent with every action.

Files opened and re-opened. The agent opened most test files at least twice, some three or four times, as it worked through failures in batches. Early in the session it also opened a few large files I would not have: a 3,000-line fixtures module and the generated type definitions for our API client. Those went into the context once and rode along for the rest of the afternoon.

Retries on the same problem. Module mocking was the hard part. Jest’s automatic hoisting of jest.mock calls behaves differently from Vitest’s vi.mock, and a dozen tests depended on that behaviour. The agent spent a long stretch trying variations: edit, run the test, read the failure, edit again. Each attempt was cheap on its own. Together they added a large block of failed attempts and error output that every later turn had to carry.

My own instructions and rules. Small per turn, but constant. Our rules file and my task description were a few thousand tokens, sent on every request.

The actual edits, the thing I was paying for in my head, were a tiny fraction of the whole.

Long sessions cost more than linearly

The part that took me longest to internalise: a session that is twice as long costs more than twice as much. Each new action re-sends everything before it, so later actions are more expensive than earlier ones. The first ten turns of my session were cheap. The last ten, carrying two hours of accumulated test output and file contents, were each many times more expensive, even though they were doing similar small edits.

This means the tail of a long session is where money quietly goes. The agent is fixing the last three failing tests, making small careful edits, and every one of those edits pays to re-read the whole afternoon.

A long paper receipt curling off the edge of a wooden desk onto the floor next to a laptop

What I changed

None of these changes made the agent less useful. Most of them made it faster too, because a smaller context is also quicker to process.

Quieter test output. This was the biggest single win. I configured the test runner to print only failures, with shortened stack traces, and added a note to the project rules telling the agent to run individual test files while iterating rather than the full suite. A full run that used to produce tens of thousands of tokens of output now produces a few hundred when most tests pass. Same information that matters, a fraction of the context.

Ignore files for the obvious heavyweights. Generated API types, large fixtures, lockfiles and build output went into the tool’s ignore configuration. The agent can still read them if I point it there, but it no longer opens them while exploring.

Split long tasks into sessions. For the next large migration, a similar one on another project, I broke the work into stages: first convert configuration and shared helpers, then commit; then convert tests directory by directory, one fresh session per directory, each starting from the committed state and a short description of what had already been done. Each session carried only what it needed. Total token use was roughly a third of the Jest migration for a similar-sized suite.

Step in on retry loops earlier. When the agent tries the same kind of fix three times, I now intervene. Partly because it is usually stuck, and partly because every failed attempt becomes permanent baggage for the rest of the session. On the mocking problem, five minutes of me reading the Vitest docs and telling the agent exactly how to restructure the mocks would have saved a big slice of that afternoon’s spend.

Keep sessions warm or end them. Knowing that caches expire, I stopped leaving long sessions idle in the middle of a task and then resuming them. If I need to step away for more than a few minutes, I either let the agent finish a coherent chunk first, or I end the session and start a fresh one with a short summary when I return. A fresh session with a two-paragraph summary is much cheaper than resuming a stale one with two hours of history.

What to look at on your own usage page

Most tools now show usage per session or per request, and the breakdown is worth ten minutes of attention. The things I check:

  • The ratio of input to output tokens. For agent work, input dwarfing output by fifty or a hundred to one is normal. If yours is several hundred to one, something is bloating the context: usually command output or large files.
  • Which sessions are outliers. In my case, a handful of long sessions accounted for most of a month’s spend. Everyday small tasks, fix this bug, add this field, were cheap.
  • Cached versus uncached input. A high share of uncached input on a long session suggests you are resuming after breaks or changing something early in the context that invalidates the cache.
  • Model choice per session. If your tool lets you pick a model, check whether the long exploratory sessions ran on the most expensive one. Reading files and running tests rarely needs the strongest model available.

Is it worth it anyway?

Honestly, yes, for that Jest migration. The session cost me somewhere around a nice lunch, and the migration would have taken me most of a weekend by hand. Even at the unoptimised rate, the maths favoured the agent.

But I had been assuming that cost scaled with output, and that assumption would have been expensive at team scale. If ten engineers each run a few long, noisy sessions a week, the difference between a tidy workflow and a sloppy one is not a rounding error. It is the difference between a tool budget nobody questions and one finance asks about.

The better frame, for me, is that a coding agent is billed like a very well-read contractor who charges by the page read, rereads the whole project folder before every sentence they write, and never skims. You would not hand that contractor a 5,000-line log file when a ten-line error would do. You would not keep them on one endless call when two shorter ones would cover it. And you would step in when they started trying the same fix for the fourth time.

My test output is quiet now, my long tasks are staged, and the usage page has become boring. Boring is the goal. The agent still reads a lot. It just no longer rereads my entire afternoon every time it changes a line.

More articles for you