Switching Models Mid-Task: What Breaks When the Cheap Model Hands Off to the Strong One

James Okonkwo

James Okonkwo

September 30, 2026

Switching Models Mid-Task: What Breaks When the Cheap Model Hands Off to the Strong One

The plan looked efficient. We needed to partition our order_events table in Postgres by month; it had grown past 400 million rows and the nightly cleanup job was taking hours. I used a fast, cheap model in my coding agent for the exploration phase: find every query that touches the table, list the foreign keys, identify the ORM models and raw SQL that would need changing. Cheap models are good at reading lots of files quickly. Then, for the actual migration and code changes, I switched the same conversation to the strongest model available. Scout with the cheap one, build with the expensive one.

The strong model wrote a careful, well-structured migration. It created the partitioned parent table, attached monthly partitions, copied data in batches, and swapped names in a transaction. It was the kind of migration I would have been proud to write. It also failed on the staging database, because order_events was referenced by a foreign key from refund_requests, which the migration did not handle. Postgres will not let you point a foreign key at a partitioned table in the way that constraint was defined without changes on both sides.

The cheap model’s exploration summary, sitting right there in the conversation, said: “No other tables have foreign keys referencing order_events.” It had searched the ORM models, found none, and concluded. The foreign key was defined in a raw SQL migration from 2021 and never mirrored in the ORM. The strong model never checked. It treated the summary as a fact, because in its context, it was one.

I have made this kind of handoff many times since, and I still do it. But I no longer assume the second model will catch what the first one got wrong. Here is what actually breaks when you switch models in the middle of a task.

The strong model inherits the weak model’s conclusions

This is the central problem, and it runs against intuition. You would expect a stronger model to spot errors in a weaker model’s work. Sometimes it does. More often, when a claim is stated plainly in the conversation history, the new model treats it as established context rather than something to verify.

From the model’s point of view, there is no difference between “a fact I was told” and “a fact another model guessed.” The conversation does not label which statements were verified by running a query and which were inferred from a partial search. The exploration summary read like a confident, accurate report. So the strong model built on it.

When I work with a human colleague, I calibrate: I know that a junior engineer’s “there are no other references” deserves a quick double-check, and a principal engineer’s usually does not. The model switch erases that calibration. The strong model is effectively reading notes without knowing who wrote them.

A backend engineer at a desk sketching database table boxes connected by lines on a large sheet of paper

Shallow exploration looks exactly like thorough exploration

The cheap model’s search was not careless in an obvious way. It opened the models directory, grepped for the table name in Python files, and reported what it found. What it did not do was search the migrations directory’s raw SQL, or query the database’s catalog for constraints. A stronger model might have. A senior engineer definitely would have run d+ order_events against the staging database, which lists every referencing foreign key in one command.

The output of a shallow search and a thorough one is the same shape: a list of findings and a confident summary. Nothing in the text tells you how hard the model looked. When the strong model picks up that summary, it gets the conclusion without the method, and it has no reason to suspect the method was thin.

Other things that break on a switch

The inherited-conclusions problem is the expensive one, but I have hit a few others often enough to plan for them.

Plans drift in the other direction. When I go the opposite way, planning with a strong model and implementing with a cheap one, the failure flips. The strong model writes a good plan with subtle constraints: “copy in batches of 50,000 keyed on id, not created_at, because created_at is not unique.” The cheap model follows the headline steps and quietly drops the subtle ones. The batching is there. The choice of key is not.

Style and conventions shift mid-diff. Different models have different habits. One names things verbosely, another tersely. One writes defensive checks everywhere, another trusts types. A diff that switches model halfway through can have two visibly different styles in adjacent functions, which is a small thing until someone has to maintain it.

Tool behaviour can differ. In some agent setups, models use tools differently: one prefers reading whole files, another reads small ranges; one runs tests eagerly, another waits to be asked. Switching mid-task can change how the agent works, not just how well it reasons, and the new pattern may not fit what the previous model set up.

The cost saving may be smaller than you think. Switching models usually means the new model processes the whole conversation from scratch, since caches are generally tied to the model. If the exploration phase built up a large context, the first request to the expensive model pays full price to read all of it. On a long exploration, that single request can eat a noticeable share of what you saved by exploring cheaply.

What I do differently now

I still split work between models. The economics are too good to ignore on large tasks, and cheap models are honestly fine at a lot of reading and searching. But I changed how the handoff happens.

The handoff is a document, not a conversation. Instead of switching models inside one long conversation, I ask the first model to write its findings into a short file, then start a fresh session with the strong model and that file. This sounds like ceremony. It has two big benefits. The strong model starts with a clean context focused on the findings rather than the noise of exploration. And I read the findings before the handoff, which is where I would have spotted “no other tables have foreign keys” as a claim worth checking.

Findings say how they were found. I ask the exploration model to include its method for each finding: “Searched app/models/ for ForeignKey to OrderEvent; none found. Did not check raw SQL migrations or database catalog.” That second sentence is the valuable one. It tells the strong model, and me, exactly where the gaps are.

The strong model verifies load-bearing claims. The first instruction in the implementation session is now something like: “Before writing the migration, verify each finding below that the migration depends on. For database structure, check the actual schema, not just the ORM.” For the partitioning task, that would have meant running a catalog query for referencing constraints, which takes one command and would have surfaced the refund table immediately.

Switch at boundaries, never mid-edit. I switch models between phases: after exploration and before planning, or after a plan is approved and before implementation. I do not switch halfway through a set of related edits. The style drift and half-understood intermediate state are not worth it.

Two coworkers at a desk, one handing a folded note to the other, who looks doubtful

Where each model belongs in my workflow

After a few months of this, my split looks roughly like the following.

The cheap model does broad, low-stakes reading: listing files that mention a symbol, summarising what a module does, finding every call site of a function, drafting a first pass of test cases from existing examples. Its mistakes here are usually omissions, and omissions are easy to check if the method is stated.

The strong model does anything where being subtly wrong is expensive: schema changes, concurrency, anything touching money or permissions, and the final plan for a multi-step change. It also gets the verification job at the start of implementation. That is the part I used to skip, and it is the part that would have saved me a failed staging deploy and an evening of untangling a half-applied migration.

For small tasks, I do not switch at all. The overhead of a handoff document is not worth it for a twenty-minute fix. I pick one model and stay with it.

The foreign key, and the lesson

We fixed the partitioning migration the next day. The refund table’s foreign key had to be dropped and replaced with a constraint that included the partition key, which meant adding the order’s creation month to refund_requests and backfilling it. It added a day to the project, and it would have been in the plan from the start if anyone, human or model, had queried the catalog.

The lesson I took away is not “cheap models are bad” or “always use the strongest model.” It is that a model switch is a handoff between two workers who cannot talk to each other, and the second one assumes the first one’s notes are accurate. That is exactly how handoffs between people go wrong too, and it has the same fix: write down what you checked and how, and have the receiver verify anything they are about to build on.

The order_events table is partitioned now. The nightly cleanup drops a partition instead of deleting millions of rows and finishes in seconds. And the exploration notes for every migration I hand off now end with a line that lists what was not checked. It is usually the most useful line in the document.

More articles for you