A Fast Cheap Model vs a Slow Strong Model for the Same Coding Task

Quinn Reed

Quinn Reed

September 30, 2026

A Fast Cheap Model vs a Slow Strong Model for the Same Coding Task

For about three months I defaulted to the fast model in my coding agent for almost everything. It answered in seconds instead of a minute, cost a fraction per session, and felt snappier to work with. The strong model was for “hard problems,” a category I defined, in practice, as problems where the fast model had already failed twice.

Then a colleague asked me a simple question in a code review: how long did this take you? The pull request was a fix for a timezone bug in our scheduling service. I had to admit the honest answer was most of an afternoon, of which the agent’s actual thinking time was maybe ten minutes. The rest was me reading diffs, spotting problems, re-prompting, and reading again. The fast model had been fast at producing wrong answers.

So I ran a small, unscientific experiment. I took three real tasks from our backlog, of different kinds, and did each one twice on separate branches: once with the fast, cheap model and once with the slow, strong one. Same repository state, same task description, same rules file. I tracked wall-clock time from first prompt to a diff I was willing to merge, how many correction rounds each needed, and roughly what each session cost.

Three tasks is not a benchmark. But the results changed my defaults, and the pattern has held up in the months since.

Task one: add cursor pagination to an endpoint

Our /api/bookings endpoint returned everything, and one customer now had enough bookings to make that painful. The task: add cursor-based pagination using the existing helper we already use on three other endpoints, update the response schema, and add tests.

Fast model: first diff in about ninety seconds. It found the pagination helper, used it correctly, updated the schema, and wrote tests modelled on the other paginated endpoints. One small issue: it forgot to add the index hint we use for the cursor column, which I pointed out, and it fixed immediately. Total time to mergeable: about fifteen minutes, most of which was me reviewing.

Strong model: first diff in about four minutes. Essentially the same change, with the index hint included, plus a slightly more thorough test for the empty-page case. Total time to mergeable: about fourteen minutes.

Verdict: a wash on time, and the fast model cost a fraction as much. This is the kind of task where there is a clear pattern in the repository to copy, and copying a pattern well does not need a strong model.

A programmer writing numbers in a paper notebook, comparing columns, with a laptop open beside a mug of tea

Task two: the timezone bug, again

I reused the bug that had prompted the experiment, resetting the branch to before my fix. Recurring appointments created by users in timezones with daylight saving were shifting by an hour after the clocks changed. The cause, which I already knew, was that we stored the recurrence’s start as a UTC timestamp and generated future occurrences by adding fixed intervals in UTC, instead of generating them in the user’s local timezone and converting each one.

Fast model: first diff in about two minutes. It identified the right function and proposed a fix: convert to the user’s timezone before adding the interval. It looked right. The tests it wrote passed. But the fix converted back to UTC before storing each occurrence using the offset of the original start date rather than each occurrence’s own date, which reintroduces the bug in a subtler form. I caught it on review, explained, and it tried again. The second attempt handled the conversion correctly but broke the case where a recurring appointment falls in the non-existent hour during the spring-forward transition. Third round got it right after I described the edge case explicitly. Total time: about seventy minutes, of which the model’s generation was a few minutes and the rest was me reading and explaining.

Strong model: first diff in about six minutes. It generated each occurrence in local time, converted each individually, and, without being asked, added handling for both the skipped hour and the repeated hour in autumn, with a comment explaining the choice it made for each. It also wrote tests that pinned specific dates around a DST transition in two timezones. I had one question about the repeated-hour behaviour, which it answered with a reasonable justification. Total time to mergeable: about twenty-five minutes.

Verdict: the strong model was slower to the first diff and much faster to a correct one. The fast model’s cost per session was lower, but I spent forty-five extra minutes of my own time on review and re-explanation. My time is not free, and it was the dominant cost.

Task three: tighten an authorisation check

A customer’s team members could see bookings from other locations within the same organisation, when they should only see their own location’s bookings unless they had an admin role. The fix involved the permission check in a service layer and a query filter, plus making sure three other endpoints that used the same service did not regress.

Fast model: first diff in about two minutes. It fixed the query filter in the main endpoint correctly. It did not notice that two of the other three endpoints called a different method on the same service that had the same flaw. The diff passed all tests, because none of the tests covered cross-location access. I only found the gap because I knew the service well enough to check its other methods. One more round fixed it.

Strong model: first diff in about five minutes. It fixed the main endpoint, then said, roughly, “the same filter is missing from list_for_org and export_for_org, which are used by the reports and export endpoints; I’ve applied the same fix there.” It added tests for cross-location access on all three.

Verdict: the fast model’s diff was correct in what it touched and incomplete in what it did not. That is the most dangerous kind of wrong, because it passes review unless the reviewer already knows where to look. On a security-relevant change, that alone decides it for me.

A sprinter and a long-distance runner on an empty athletics track at sunrise, seen from behind

What the numbers said

Adding up the three tasks: the fast model sessions together cost well under a third of the strong model sessions. The fast model produced its first diffs roughly three times quicker. And the fast model’s total time to mergeable, across all three tasks, was about an hour longer, almost all of it spent on my review and re-prompting on tasks two and three.

The session cost difference was a few dollars. The time difference was an hour of senior engineering time. It was not close.

But that total hides the real pattern. On task one, the fast model was the right choice: same quality, lower cost, no extra review. On tasks two and three, it was the wrong choice by a wide margin. The question is not which model is better. It is how to tell, before you start, which kind of task you have.

How I choose now

The distinction that best predicts the outcome, for me, is whether the task has a correct pattern already in the repository or requires reasoning the repository does not already contain.

Fast model when the pattern exists. Adding pagination like the other endpoints. Writing tests in the style of neighbouring tests. Renames, straightforward refactors, boilerplate, updating call sites after a signature change. Explaining what a module does. In these tasks the answer is mostly “do what the codebase already does here,” and a fast model is excellent at that.

Strong model when the edge cases are the task. Time and timezones. Concurrency. Money and rounding. Permissions and authorisation. Migrations with data. Anything where the obvious fix is often subtly wrong, and where a correct answer requires thinking about cases that do not appear in the code yet. The strong model’s extra minutes go into exactly the reasoning I would otherwise have to supply in review.

Strong model when I cannot easily verify. If I know the code well, I can catch a fast model’s mistakes quickly, and the cheap first draft may be fine. If I am in an unfamiliar part of the codebase, I am a worse reviewer, and I want the model that is more likely to be right on its own. Task three would have gone badly if a colleague less familiar with the service had been reviewing.

People sometimes suggest using both: scout with the cheap model and build with the strong one. I do that on large tasks, but it has its own trap, because the strong model tends to trust what the cheap model handed off mid-task without re-checking. For the kind of single, contained task I tested here, picking one model up front was simpler and worked better.

The metric that changed my mind

The fast model feels faster because the thing you notice is time to first response. You type, and something appears almost immediately. The strong model’s long pause feels like waste.

But the time that matters is time to merged, including your review and every correction round. On tasks that need real reasoning, the fast model’s quick first draft is not a head start. It is a draft you have to debug, and debugging someone else’s subtle mistake takes longer than reviewing a correct change.

I still use the fast model most days, for most small tasks. What changed is the default for anything involving time, money, permissions or data migrations: I start with the strong model and accept the wait. The four-minute pause while it thinks is much cheaper than the forty-five minutes I would spend explaining daylight saving time to a model that answered in ninety seconds.

More articles for you