What Happens When a Local LLM Doesn’t Fit in VRAM?
Marcus Feldmann
September 27, 2026
The card in my test bench this summer was an RTX 4070 with 12 GB of video memory, and the model I wanted to run was a 32B at Q4, a file of about 19.9 GB. I knew it would not fit. I wanted to see exactly how it would not fit, because “doesn’t fit in VRAM” gets described online as if it were one thing, and in practice it is at least four different things depending on which software you use and how you configured it.
Over one long evening the same model on the same card crashed outright, ran at a respectable-sounding “partially offloaded” speed that was actually miserable, ran even slower while claiming to be fully on the GPU, and finally ran at a genuinely usable speed once I stopped fighting the hardware. This is what happens in each case, why, and what you can actually do about it on a Windows or Linux PC with an NVIDIA card.
First, what is competing for the card’s memory
Video memory has to hold three things for a model to run entirely on the GPU:
- The weights. For my 32B at Q4, about 19.9 GB. For a 14B at Q4, about 9 GB.
- The KV cache, the model’s working memory of the conversation so far, which grows with the context window you set. For a 32B-class model it is roughly a quarter of a megabyte per token at default precision, so an 8,000-token window costs about 2 GB and a 32,000-token window about 8 GB.
- Overhead: the CUDA context, scratch buffers, and whatever else is already using the card. On a Windows desktop with a browser open and two monitors attached, that “whatever else” was 1.3 GB on my machine before I loaded anything.
If you are unsure whether a model could ever fit, do that sum before downloading. For me it was obvious: nearly 20 GB of weights into roughly 10.5 GB of genuinely free memory was never going to happen. The question was what the software would do about it.
Outcome 1: it crashes
I started with llama.cpp’s server and asked it to put every layer on the GPU (-ngl 99). It began allocating, ran out of video memory partway through loading, and stopped with a CUDA out-of-memory error. No chat window, no slow answer, just a failure message and a return to the command prompt.
This is the most honest outcome. You asked for something impossible and were told so. The same crash can also happen later rather than at load time: a model that fits with a small context can run out of memory when you raise the context window or when a long conversation pushes the cache past what was reserved, depending on how the runtime allocates. If a model loads fine and then dies when you paste a long document, that is almost always the cache.
Outcome 2: it splits the model between GPU and CPU
Next I loaded the same file in Ollama. It did not crash. Ollama estimates how much video memory is available, puts as many of the model’s layers on the GPU as fit, and runs the rest on the CPU from system RAM. ollama ps showed the split: roughly half on the GPU, half on the CPU.
“Half on the GPU” sounds like “half the speed.” It is much worse than that, and the reason is worth understanding because it explains almost everything about spills.
When a model generates a token, it has to read all of its weights once. The GPU part reads from video memory at around 500 GB/s on a 4070. The CPU part reads from system memory, which on my dual-channel DDR5 delivers maybe 65 GB/s in practice. So for each token, the roughly 10 GB on the GPU takes about 0.02 seconds to read, and the roughly 10 GB on the CPU takes about 0.15 seconds. The slow half dominates completely.
The result: about 6 tokens per second. For comparison, a 14B model that fits entirely in the 4070’s memory runs at around 40 tokens per second on the same machine. The split did not halve my speed compared with a model that fits. It cut it to a sixth, while the GPU sat mostly idle waiting for the CPU to finish its layers.

The rule of thumb that falls out of this: once even a modest fraction of the model lives in system RAM, your speed is set by system RAM. Offloading 10 percent of the layers costs you noticeably. Offloading half means you are effectively running a CPU model with a GPU helping out. Prompt processing, the step where the model reads your input, holds up better because the GPU still does much of the heavy computation, but generation speed falls hard.
Partial offload is still useful. Six tokens per second is readable, and for a question where you want the bigger model’s answer and are happy to wait, it works. It just is not what the phrase “GPU-accelerated” makes you picture.
Outcome 3: the driver spills silently, and it is the worst of all
Then I tried LM Studio on Windows, set GPU offload to maximum, and something strange happened. It loaded. The interface suggested the whole model was on the GPU. And it ran at about 2 to 3 tokens per second, slower than Ollama’s honest half-and-half split.
The cause is a feature in NVIDIA’s Windows driver called System Memory Fallback, introduced in 2023. When an application asks for more video memory than exists, instead of returning an out-of-memory error, the driver quietly starts using system RAM as overflow and shuttles data back and forth across the PCIe bus. It was added so that games and creative apps crash less often. For language models it is a trap: the “overflow” data has to cross PCIe, which on my board manages around 25 GB/s in practice, even slower than the CPU reading its own RAM. And because the application thinks it has all the memory it asked for, it never falls back to a sensible CPU split.
The giveaway is in Task Manager. Under the GPU tab, “Dedicated GPU memory” is pinned at the maximum and “Shared GPU memory” is climbing by several gigabytes. If you see that while a model runs slowly, the driver is spilling.
You can turn it off. In the NVIDIA Control Panel, under Manage 3D Settings, there is a “CUDA – Sysmem Fallback Policy” option. Set it to “Prefer No Sysmem Fallback,” either globally or just for your inference program. Now an oversized model crashes or, in runtimes that handle it, gets split between GPU and CPU properly instead. I much prefer a clear error to a mystery slowdown. Linux drivers do not have this behaviour, which is one reason people report such different results for the same card on the two systems.

Outcome 4: it fits at 4K and spills at 32K
The last trap is quieter. Take a model that nearly fits, like a 14B at Q4 on a 12 GB card. With a 4,000-token context it loads entirely on the GPU and flies. Raise the context window to 32,000 tokens so you can paste long documents, and the cache grows by several gigabytes. Ollama recalculates, finds that the model plus the bigger cache no longer fits, and moves some layers to the CPU. Nothing tells you this happened except that everything got slower.
If a model’s speed changes dramatically when you change context settings, check the GPU/CPU split again. The model did not change. The memory budget did.
What actually helps
Here is what I tried after the four failures, roughly in order of how much each one bought me.
Free the memory you are already wasting
Close the browser or turn off its hardware acceleration, quit anything else that uses the GPU, and if you have a second display output on the motherboard, plug a monitor into that instead. On my machine that recovered about a gigabyte. Sometimes a gigabyte is the difference between one layer on the CPU and none.
Quantize the cache
llama.cpp and Ollama can store the KV cache at 8-bit instead of 16-bit, which roughly halves its size with very little quality cost. In llama.cpp that means enabling flash attention and setting the cache type to q8_0; in Ollama, the OLLAMA_FLASH_ATTENTION and OLLAMA_KV_CACHE_TYPE environment variables. For long-context work on a card that is nearly big enough, this is the single most effective setting.
Use a smaller quantization, carefully
Dropping from Q4_K_M to a 3-bit quant shrinks the weights by about a fifth. My 32B at 3 bits was still too big for 12 GB, but a 24B model at a 3-bit quant fit with a short context. Quality falls off faster below 4 bits, so test it on your own prompts. In my testing, a 14B at Q4 that fits entirely was usually a better experience than a 32B squeezed to 3 bits and split across the CPU.
Pick a model built for splitting
This was the real win of the evening. Mixture-of-experts models store many parameters but use only a small subset of “experts” for each token. llama.cpp lets you keep the parts that every token uses (the attention layers and shared weights) on the GPU and push the experts to system RAM, with the --n-cpu-moe option or a tensor override pattern. Because each token only touches a few experts, the CPU side reads far less data per token than in a normal split.
Qwen3 30B-A3B at Q4 is about 18.6 GB, too big for my card. Loaded this way, with attention on the 4070 and the experts in DDR5, it ran at around 30 tokens per second, five times faster than the dense 32B split, with answers that were close in quality for my everyday prompts. OpenAI’s gpt-oss-20b, at about 13 GB, only needed a few expert layers moved to the CPU and ran faster still.
Or accept the obvious
If you regularly want dense 32B models at full speed, you need a card with 24 GB, or two cards that add up to it; llama.cpp can split layers across multiple GPUs. No setting turns 12 GB into 20.
The short version
When a model does not fit in video memory, one of four things happens: the runtime crashes with an out-of-memory error; it splits the model between GPU and CPU and runs at the speed of your system RAM; on Windows, the NVIDIA driver silently spills to system memory over PCIe and runs slower still; or a model that fitted with a short context starts spilling when you raise it. The crash is the honest one. The silent spill is the one to hunt down and switch off. And if you are stuck on a 12 GB card, a mixture-of-experts model with its experts offloaded to the CPU is the best way to get a big model’s answers at a small model’s speed.