Caching Prompt Context vs Paying to Resend the Same Files
Casey Holt
September 30, 2026
The number that got my attention was not the bill. It was a single column in our usage export: cache read tokens. For the internal review agent my team runs against pull requests, that column was close to zero. Not low. Close to zero. Meanwhile the cache write column was enormous, and cache writes are billed at a premium over ordinary input.
We had turned on prompt caching months earlier, felt good about it, and never checked. The agent had been paying extra, on every single call, to store a prefix it would never read back. In plain terms, turning caching on had made it more expensive than leaving it off.
This is the story of how that happened, what we changed, and the rule I now use to decide whether a given agent loop should be built around a cache at all or should simply pay to resend the same files every time.
What the cache actually matches on
Most people picture prompt caching as the provider noticing that you sent the same files again and giving you a discount. That is roughly the effect, but the mechanism is much stricter, and the strictness is where every failure I have seen comes from.
The cache works on the prefix of the request. The provider takes the start of your prompt, everything up to a point, and stores the model’s internal work for it. On the next request, if the start is identical down to the byte, that stored work is reused and billed at a small fraction of the normal input price. The moment one token differs, everything from that token onward is treated as new. It does not matter that the other ninety-nine percent of the files are unchanged if they come after the change.
Some providers make you mark where the cacheable section ends; others cache automatically once the prefix passes a minimum length, usually around a thousand tokens. The details vary and they keep shifting, so check the current pricing page rather than trusting anything I write about exact multipliers. The shape has been stable for a while, though: writing to the cache costs somewhat more than normal input, reading from it costs far less, and the entry expires after a few minutes without use unless you pay for a longer lifetime.
Those three facts are the whole game. Everything else is arithmetic.

Three things that broke our prefix
Once I knew to look, I diffed two consecutive requests from the review agent. They were supposed to share a long, stable opening: the system instructions, the tool definitions, our repository conventions file, and a map of the codebase. The diff showed changes in all of it.
A timestamp on line two. Someone, possibly me, had added “Current date and time:” near the top of the system prompt so the agent could reason about how recent a commit was. It was a sensible idea placed in the worst possible spot. Because it changed on every call, the cache matched nothing after the first line, and every request wrote the entire prefix to the cache again.
Tool definitions in unstable order. Our agent framework assembled its tool list from a dictionary that was populated as plugins loaded. Load order depended on which plugin finished initialising first, which varied between runs. The same eleven tools, described with the same words, arrived in a different order often enough to break the match on a large share of calls. Nobody would ever spot that by reading a prompt, because both versions look identical to a human.
A repository map that included file sizes. The codebase map listed each file with its size and last-modified time, which the agent never used. Every commit changed some of those numbers, so the map, sitting right in the middle of the “stable” section, was different on almost every pull request. Everything after it, including the conventions file, was re-sent at full price or worse.
None of these were bugs in the usual sense. Each one was a small, reasonable decision by someone who was thinking about what the model should see, not about what the cache needed to stay identical.
The arithmetic that makes a broken cache worse than none
It is worth doing the sums once, because they explain why “just turn caching on” is not free.
Say the stable part of the prompt is 60,000 tokens and a review takes 40 model calls. Use rough, illustrative multipliers: writing to the cache costs 1.25 times normal input, reading costs 0.1 times.
- No caching: 40 calls × 60,000 tokens at full price is 2.4 million token-equivalents.
- Caching that works: one write at 1.25× (75,000) plus 39 reads at 0.1× (234,000) is about 309,000. Roughly an eighth of the uncached cost.
- Caching that breaks every call: 40 writes at 1.25× is 3 million. A quarter more than doing nothing.
That last line is where we were. The feature we had switched on for savings was a steady surcharge, and because the output looked fine and the agent was fast, nothing prompted anyone to look.
It also shows the opposite point, which matters more for deciding when to bother. The saving comes almost entirely from reads. A cache entry that is written once and read zero or one times is a loss or a wash. The value depends on how many times the same prefix gets reused before it expires, and that is a property of the workload, not a setting.
What we changed
The fix was mostly reordering. I started thinking of the prompt as layers, sorted by how often each one changes, with the slowest-changing material first.
- Frozen: system instructions and tool definitions. The tools are now sorted by name before serialisation, and the instructions contain nothing that varies at runtime. This layer changes when we deploy a new version of the agent, and at no other time.
- Slow: the conventions file and a trimmed codebase map with paths only, no sizes or dates. This changes when the repository structure changes, which is a few times a week.
- Per task: the pull request description and the diff under review. This changes for every review, but stays constant across the 30 to 50 calls within one review.
- Volatile: the date, the current step, tool results and the running conversation. This goes last, and nobody tries to cache it.
The timestamp still exists. It just lives at the bottom now, where changing it costs nothing. We put cache breakpoints at the end of the first three layers, so a change to the diff invalidates only the diff layer and what follows, not the instructions above it.
After the change, cache reads became the largest input line on every review, and the cost per review fell to roughly a fifth of what it had been. The model’s output did not change in any way we could measure, which makes sense: it saw the same material, just in a different order.
One caution about reordering: models do pay attention to position, and material near the end of a long prompt tends to weigh more heavily. Moving the diff under review below the static files suited us, because the diff is what the agent should focus on. If your most important instruction lives in the frozen layer, repeat a short version of it near the end rather than moving the whole block and breaking the prefix again.

When paying to resend is the right call
Caching is not always the answer, and I have talked two teams out of building around it. The question I ask is simple: how many times will this exact prefix be read before it expires?
Long gaps between calls. We also have a nightly job that summarises the day’s merged changes. It makes one large call per repository, once every 24 hours. No default cache lifetime survives that, and paying for a longer one to cover a single read makes no sense. That job pays full price for its context and it should. Adding cache markers would only add the write premium.
Human-paced sessions. An interactive session where a developer reads each response, thinks, and types for six minutes will miss a five-minute cache regularly. Whether it is worth paying for a longer lifetime depends on how long the pauses are and how large the prefix is. For small prefixes, I leave it alone.
Small prefixes. Below the provider’s minimum cacheable length nothing is cached anyway, and just above it the absolute savings are tiny. If the whole stable context is a few thousand tokens, the engineering effort of keeping it byte-identical is better spent elsewhere.
Context that genuinely changes every turn. Some agent designs rebuild the file set on every step, pulling in whatever files look relevant to the current sub-task. If the selected files differ each time, there is no stable prefix to cache. You can restructure so the core files stay pinned at the top and only the extras rotate below them, but if the design depends on a fully fresh selection, accept the resend cost and budget for it.
The inverse is the case where caching earns its keep: a tight loop, many calls a minute, with a large block of the same files at the start of every request. That describes most coding agents in the middle of a task, which is why it is worth getting right there.
Caching only changes the price of re-reading; it does not change how much gets re-read. If the per-call context keeps climbing as the task goes on, the cache discounts each climb but the total still grows. A long agent task accumulates that much context because every step re-sends the files and history it has already read, and a cache does nothing about that. The cache is the discount, not the diet.
How to tell whether yours is working
The fastest check is one I wish we had run on day one. Every major API returns cache figures in the usage section of each response: tokens written to the cache, tokens read from it, and uncached input. Log those three numbers per call, per job.
Then look at the ratio over a full task. In a healthy agent loop, cache reads should dominate input within the first few calls and stay dominant. The warning signs:
- Writes on every call: something near the top of the prompt is changing. Diff two consecutive raw requests and look for the first differing byte. It is usually a timestamp, a random ID, or unstable ordering.
- Reads for a while, then a sudden stretch of writes: the cache expired during a pause, or an early file in the prefix was edited. If it is the agent editing a file that sits in the cached section, consider moving frequently edited files lower in the layers.
- Zero cache figures at all: the prefix is below the minimum, or caching is not enabled for the model or endpoint you are using.
I also added a crude alarm: if a job’s cache write tokens exceed its cache read tokens over a day, it posts a message to our team channel. It has fired twice since, once for a new plugin that added a request ID to its tool description and once for a teammate who put a build number in the system prompt. Both fixes were one-line changes, and both would have run for months otherwise.
The rule I use now
If an agent loop sends the same large block of files more than a handful of times within a few minutes, I build it around the cache on purpose: stable material first, sorted and scrubbed of anything that varies, volatile material last, and cache figures logged from the first run. If a job reads its context once, or with long pauses, I skip the cache markers and pay full price, because the write premium buys nothing.
What I no longer do is switch caching on and assume it is working. The feature is only as good as the prefix you give it, and a prefix that looks stable to a human can differ on every call because of one line nobody thought about.