Pinning the Model vs Auto: Why Yesterday’s Working Prompt Fails Today
Dmitri Voronov
September 30, 2026
The client called on a Tuesday morning, polite and slightly amused. “Dmitri, we’re having a sale I didn’t know about. The three-seater in oak is on the website for one euro thirty.”
The client is a family furniture business with two showrooms and an online shop I maintain on contract. Every week their main supplier sends a spreadsheet of updated prices and stock, in a format that changes just often enough to be annoying. For about four months I had handled it with a saved prompt: I gave my coding agent the new spreadsheet and the prompt, it updated the import mapping and validation rules in the repository to match any format changes, ran the import against staging, and showed me a summary. I skimmed it, merged, and the import ran against production.
It had worked sixteen weeks in a row. On the seventeenth, the agent parsed “1.299,00” as one point two nine nine. Every price over a thousand euros came out a thousand times too small. The summary said “412 products updated”, which was true.
We fixed the prices within the hour from the previous day’s backup. Working out why it happened took the rest of the week, and the answer was not what I expected.
I could not reproduce it
The first thing I did was run the exact same prompt against the exact same spreadsheet on a clean branch. It produced the correct mapping: European number format, comma as decimal separator, prices in the thousands. I ran it again. Correct again. The third time, it produced the broken version.
That is a uniquely unpleasant kind of bug. The prompt had not changed. The spreadsheet had not changed. The repository was reset to the same commit. And the output changed from run to run.
My agent was set to Auto, where the tool picks a model for each request based on the task, current load and cost. I had left it there because it was the default and it had always worked. When I looked at the request details, the correct runs and the broken run had been served by different models. The Tuesday import had been handled by a model that, for this particular spreadsheet, did not infer the locale from the sample rows.

What Auto and pinning actually mean
It is worth being precise here, because “pinning the model” means different things in different places, and I had the vague idea that choosing a model by name was the same as pinning it.
Auto means the tool chooses. Different requests in the same week, or even the same session, may go to different models. That is the point of it: it can send easy work to cheaper, faster models and hard work to stronger ones, and it can route around a provider that is having a bad day. For interactive work it is often excellent. For a job you expect to behave identically every Tuesday, it means the one thing you assumed was constant is not.
A model name in a picker is often an alias. “The latest version of model X” is a pointer, and the provider moves it when a new version ships. You chose the family, not the exact weights. Most of the time the new version is better. Occasionally it is better on average and worse on your specific task.
A dated snapshot is the closest thing to a real pin. Most providers publish model identifiers with a version or date in them that refer to one fixed model. That model will behave the same way until it is retired, and providers announce retirement dates in advance. If you call the API directly, you can use these identifiers. Inside an editor’s agent you may or may not get that level of control, depending on the tool.
I had been on Auto. Even if I had chosen a model by name, I would only have been partly protected.
The model was not the only thing that moved
Before blaming everything on the router, I listed everything else that could make the same prompt behave differently on a different day. It was a longer list than I expected.
- The tool’s own instructions. Agent tools wrap your prompt in their own system prompt and tool definitions, and they update those with new releases. My editor had updated itself twice in those four months.
- Your rules files. I had added a line to the project’s agent instructions a few weeks earlier asking it to use the standard library’s number parsing instead of hand-written regular expressions. I had forgotten it existed. The standard parser assumes the server’s locale unless told otherwise, and the server’s locale uses a dot for decimals.
- The repository. The agent reads surrounding code for context. A refactor elsewhere can change what it sees and therefore what it writes.
- Randomness. Even a fully pinned model does not always give the same answer twice at normal settings. Usually the differences are cosmetic. Sometimes one of them is a decision.
So even with a perfect pin, “yesterday’s prompt fails today” can happen. Pinning removes the biggest source of variation. It does not remove all of them.
The real bug was in my prompt
This is the part I did not enjoy admitting. My saved prompt said, more or less: “Here is this week’s supplier spreadsheet. Update the import mapping and validation to match its format, run it against staging, and summarise the changes.”
Nowhere did it say that the supplier is in Germany, that prices use a comma as the decimal separator and a dot for thousands, or that no product in the catalogue costs less than about fifteen euros. For sixteen weeks, the models that happened to handle the job had worked out the number format from the sample rows. One week, the model that handled it made the other reasonable guess.
The prompt did not work reliably for sixteen weeks. It depended on a guess, and the guess happened to come out right sixteen times. Pinning a model would have kept that lucky guess stable for longer. It would not have made the prompt correct, and the next model migration would have broken it just the same.
So I fixed the prompt first. It now states the locale and number format explicitly, includes three sample rows with the expected parsed values, and says that any price below ten euros or above twenty thousand must be flagged, not imported. I also added a validation test to the import itself that fails if the median price moves by more than a set percentage from the previous week. That test would have stopped the Tuesday import on its own, regardless of which model wrote the mapping.

When I pin, and when I leave it on Auto
After the fix, I went through all my recurring agent jobs across clients and sorted them.
I pin for repeatable jobs that run without me reading every line. Weekly imports, report generation, scheduled content updates, the scripted parts of deployments. Anything where the output is trusted because it worked last time. For these I use a specific model version, either through the API or through the most specific setting my tools allow, and I write down which one.
I leave Auto on for interactive work. When I am sitting with the agent, building a feature, reading each diff, and correcting as I go, a change of model from one request to the next does not matter much, because I am the consistency. The router’s choice of a faster model for a small edit is a benefit, not a risk.
I pin for client-facing output that has to match a previous version. One client has product descriptions generated in a particular house style. When the style drifted slightly after a model update, they noticed before I did. Now that job is pinned, and style changes happen when I choose.
What pinning costs
Pinning is not free, and I have seen people treat it as a permanent answer.
A pinned model gets retired. When it does, you migrate on the provider’s schedule, and if you have not tested the replacement, you are back to a surprise on a Tuesday, just a later one. You also miss improvements: newer versions are often better at exactly the kind of structured extraction my import job does. And the pinned version may cost more per token than a newer, cheaper model that does the job equally well.
So each pinned job now comes with two things.
A small set of test cases. For the import job it is six real spreadsheets from past weeks, including the one with the European thousands separators and one where the supplier renamed half the columns, each with the expected mapping output. When I consider moving to a newer model, I run all six through it and compare. It takes ten minutes. It is the difference between migrating and hoping.
A calendar reminder. Every three months I check whether the pinned model has a retirement date, and whether a newer one passes the test cases. If it does, I move deliberately, on a Monday, with the previous week’s backup ready.
What I tell clients now
Most of my clients do not care which model runs their jobs. What they care about is whether the price on the website is right. After the sofa incident, the owner asked me a reasonable question: “Is the AI going to do this again?”
The honest answer was: the model might make a different guess again, so we stopped letting it guess. The prompt states the rules, the model is fixed for this job, and a check in the import stops the run if the numbers look wrong, whoever wrote the code. The model is one part of the process. It should not be the only part standing between a spreadsheet and the live shop.
The three-seater in oak is back at its proper price. The median-price check has fired once since, when the supplier really did run a clearance sale, and I got to approve it knowingly instead of finding out from a phone call.