A Local Model for Boilerplate vs a Hosted Model for the Hard Diff

Sora Nguyen

Sora Nguyen

September 30, 2026

A Local Model for Boilerplate vs a Hosted Model for the Hard Diff

The plan fit on a sticky note. The local model does the boring parts. The hosted model does the hard parts. My laptop was already running a coder model all day for small questions, so why was I paying a hosted model to write DTOs, test fixtures and the fortieth CRUD handler of my life?

I ran that split for three weeks on a side project: a small booking API in TypeScript with a Postgres database and a React admin panel. The local model was a 14B-class coder model at Q4, served through Ollama, on my 24 GB M-series MacBook. The hosted model was whichever strong model my agent was pointed at. Boilerplate went local. Anything with a real decision in it went hosted.

By the end of week three I had stopped routing work to the local model inside the agent at all. Not because it was bad at boilerplate. It was surprisingly decent at boilerplate. It lost for reasons I had not put on the sticky note, and the model that actually beat it was not the strong one.

Boilerplate inside an agent is not typing

When I pictured “boilerplate”, I pictured the output: a file of predictable code. What I forgot is everything an agent has to do to produce that file in a real repository.

Take a simple job: add a cancellations resource that follows the same shape as the existing bookings resource. Route, handler, validation schema, repository function, types, a test file. Every line is a copy of a pattern that already exists. Perfect local work, I thought.

To do it, the agent has to read four or five existing files to learn the pattern, decide where each new piece goes, emit the edits in whatever format the tool expects, run the type checker and tests, read the output, and fix what broke. The code generation is maybe a third of the job. The rest is reading, following a protocol, and reacting to tool results.

The local model was fine at the third that is code. It struggled with the other two thirds.

Where the local model actually failed

Edit format. My agent applies changes as search-and-replace blocks: find this exact text, replace it with that. The local model got the new code right and the search text slightly wrong about one time in five. A trailing comma it remembered but the file did not have. Two spaces where the file had a tab. The edit failed to apply, the agent retried, and each retry burned another full pass over the context. A hosted model with the same instructions almost never missed a search block. This is not about intelligence in the IQ-test sense. It is about precise copying under a strict format, and small quantised models are measurably sloppier at it.

Tool calls. Once or twice a session, the local model emitted a tool call with malformed arguments, or described the call in prose instead of making it. The agent handled it by asking again, but “handled it” meant another round trip, and on a laptop every round trip is slow.

Prefill. This was the one I should have predicted, because I have written about unified memory before. To follow the bookings pattern, the model needs the bookings files in context. Five files, the tool definitions and the agent’s system prompt came to roughly 18,000 tokens before the model wrote a single character. On my machine, processing that prompt took long enough that I would look at my phone. Generation itself was quick. The waiting was all at the start, and it happened again on every turn of the loop, because each turn resends the conversation.

Memory with the rest of my life open. A 14B model at Q4 fits comfortably in 24 GB on its own. With a 32k context, an editor, a browser with documentation, Docker running Postgres and the dev server, it no longer fits comfortably. On two long sessions macOS started compressing memory hard and the whole machine went soft. The model was never the only thing on the laptop, and the agent loop makes the context grow.

Workbench with a stack of identical pre-cut wooden pieces beside one shaped piece held in a vise

The handoff made the hard diff harder

The split also had a cost on the hosted side that I did not expect.

The cancellations feature was not all boilerplate. It had one genuinely tricky rule: a cancellation within 24 hours of the booking start triggers a partial refund, calculated in the venue’s timezone, and a booking that has already been partly refunded cannot be cancelled twice. That rule was the hard diff, and it went to the hosted model.

But the hosted model was now working on top of the local model’s scaffold. The scaffold used a slightly different naming convention for the repository function than the rest of the codebase, typed the refund amount as a plain number where the rest of the project used an integer-cents type, and put the validation schema in a new file instead of the shared one. None of it was wrong enough to fail tests. All of it was wrong enough that the hosted model either had to work around it or spend its first several steps cleaning it up.

It chose to clean it up, which was the right call and also meant I paid the hosted model to redo a chunk of the boilerplate I had routed away to save money. On two other features the same thing happened. The scaffold shapes the hard part. The names, the types and the file layout are decisions, even when they look like typing, and the model making the hard decision wants to make those too.

So the first rule I took away was: split by change, not by line. If the boilerplate and the hard logic land in the same pull request, one model should write both. Handing the easy 80 percent of a change to one model and the hard 20 percent to another sounds efficient and mostly produces a second model fixing the first one’s taste.

The competitor I was ignoring

The real surprise came when I tried the same boilerplate tasks with the fast, cheap hosted model my agent also offers, the one I had skipped because “local is free”.

On the cancellations scaffold, the cheap hosted model finished before the local model had finished processing its first prompt. It missed zero search blocks across the week I tried it. It needed no memory on my laptop. And the cost of a scaffold was a few cents.

That reframed the whole question. I had been comparing local against the strong model, which made local look like a bargain for easy work. The honest comparison for boilerplate inside an agent is local against the cheapest hosted model that can follow the edit format reliably. On that comparison, local lost on speed, lost on reliability, and won only on a price difference small enough that one failed edit retry wiped it out in my own time.

That is specific to my hardware. On a desktop with a big GPU and fast prompt processing, the speed picture changes. On a 16 GB laptop it gets worse. But I suspect a lot of people running local coder models on laptops are in my position without having timed it.

Where the local model kept its job

I did not uninstall anything. The local model moved out of the agent loop and into a different role, where its weaknesses do not matter.

Single-shot generation with the pattern pasted in. Turning a JSON sample into a type definition. Writing a test fixture from a schema. Generating twenty rows of realistic seed data. Drafting a regex and explaining it. I select the input, hit a hotkey, and the output lands in a scratch buffer. No tools, no edit format, no loop, a few hundred tokens of context. The local model is instant at this and never touches the repository directly, so a sloppy edit cannot happen.

Questions about code I already have open. “What does this function return when the list is empty?” The context is one file. The answer does not need to be perfect, because I am reading the code anyway.

No network. On trains, on bad hotel Wi-Fi, on a plane, the local model is the only model. I wrote plenty of honest boilerplate on a four-hour train ride in week two, pasting patterns in by hand. It was slower than the agent and much faster than nothing.

Laptop and headphones on a fold-down table by a train window with countryside passing outside

How I split the work now

The sticky note has three lines on it now instead of two.

  • Local, outside the agent: self-contained snippets where I supply the pattern and paste the result myself. Anything offline.
  • Cheap hosted model, inside the agent: changes that are pure pattern-copying and do not contain a hard rule. Renames, a new resource that mirrors an existing one exactly, updating call sites.
  • Strong hosted model, inside the agent: any change that contains a real decision, including all of the boilerplate in the same change.

The hard diff never goes local. Not because a local model cannot reason about a refund window. Sometimes it can. It is because the hard diff is exactly where I cannot afford an edit that silently applies to the wrong place, a rule it half-remembers from the scaffold, or a twelve-minute session that should have taken four.

A test before you build the split

If you are tempted by the same sticky note, run this before routing anything:

  1. Pick one real boilerplate task from your repository, something that mirrors an existing pattern.
  2. Run it five times through your agent with the local model, from a clean branch each time. Count failed edits and malformed tool calls.
  3. Time the first response with realistic context, not a toy prompt. Do it with your normal apps open.
  4. Run the same task with the cheapest hosted model your agent supports. Compare wall-clock time to a mergeable diff, not tokens.

If the local model gets through all five runs cleanly and responds in a time you can live with, you have better hardware or a better-suited model than I do, and the split may work for you. If it misses even one edit in five, it will cost you more in retries and attention than it saves, and it belongs next to the agent rather than inside it.

The local model is a good tool. It turned out to be a good tool for a job next to the agent, not a job inside it. Letting it write boilerplate that the strong model then had to redo was the most expensive way I have found to save money.

More articles for you