How Much RAM Do You Actually Need to Run an LLM Locally?
Marcus Webb
September 27, 2026
The first local model I tried to run properly was a 13B chat model on a laptop with 16 GB of RAM. The download page said the file was 7.9 GB. I did the arithmetic every beginner does: 16 minus 8 is 8, so I had room to spare. Twenty minutes later I was watching the fan spin, the cursor stutter, and the model produce one word every four seconds while the disk light blinked like it was trying to signal a ship.
Nothing was broken. I had simply counted one number and ignored three others. The file size is the smallest part of the memory bill, and the part that grows while you use the model is the part nobody puts on the download page.
This piece is the chart I wish I had printed and taped to the monitor: which class of model fits in how much memory, what the extra overhead actually is, and where the chart stops being honest. It is deliberately not a shopping guide. Whether you should spend the money on more system RAM or on a graphics card is a different decision with different trade-offs. Here the only question is the size of the box the model has to fit into.
The one formula that does most of the work
A language model is mostly a very large list of numbers called weights. How much memory those weights need depends on two things: how many of them there are (the parameter count, the “8B” or “70B” in the name) and how many bits each one is stored in.
Almost nobody running models at home uses the full 16-bit versions. You use a quantized copy, where each weight is squeezed into fewer bits. The common GGUF files you get through Ollama, LM Studio or llama.cpp are labelled with codes like Q4_K_M or Q8_0. The practical translation:
- FP16 / BF16: 2 bytes per parameter. An 8B model is about 16 GB.
- Q8_0: a little over 1 byte per parameter. An 8B model is about 8.5 GB.
- Q4_K_M: roughly 0.6 bytes per parameter once you include the bookkeeping. An 8B model is about 4.9 GB.
So the rough rule is: parameters in billions × 0.6 = gigabytes for a Q4 file. That rule is good enough to reject a model in five seconds. It is not good enough to promise a model will fit, because the file is only the first line of the bill.
The three costs the file size leaves out
1. The context, which grows while you talk
When a model reads your prompt and writes its reply, it keeps a running memory of every token so far. That is the KV cache, and it is the reason my 13B was fine for the first question and miserable by the fifth.
The cache cost per token is fixed by the model’s architecture. For Llama 3.1 8B at default precision it works out to about 128 KB per token. That sounds tiny until you multiply it:
- 4,000 tokens of context: about 0.5 GB
- 8,000 tokens: about 1 GB
- 32,000 tokens: about 4 GB
- 128,000 tokens: about 16 GB, more than three times the model itself
A 70B model in the same family needs roughly 320 KB per token, so 32,000 tokens of context costs about 10 GB on top of the weights. That is why “supports 128K context” on a model card and “runs 128K context on your machine” are different sentences. Most runtimes default to a modest window for exactly this reason, and many chat front ends quietly raise it. If memory suddenly climbs after you paste a long document, this is where it went.

2. The runtime and its scratch space
The inference engine itself needs working memory for intermediate results as it pushes each token through the layers. On a GPU there is also the driver and the CUDA context. Budget somewhere between 0.5 and 1.5 GB for this, more for bigger models and larger batch sizes. It is not dramatic, but it is exactly the half gigabyte that turns “just fits” into “does not.”
3. Everything else on the machine
This is the one I got wrong. The operating system, a browser with a dozen tabs, a chat app built on Electron, maybe Docker: on a normal desktop that is easily 4 to 8 GB before a model loads. On my 16 GB laptop I had perhaps 9 GB genuinely free. A 7.9 GB model file plus a couple of gigabytes of context plus runtime overhead was never going to live there. The system did what systems do and started paging to disk, and a model that is paging is not slow in the “a bit sluggish” sense. It is slow in the “go make dinner” sense.
The chart
Here is the version I actually use. “Comfortable” means the model loads at Q4_K_M with around 8,000 tokens of context, the runtime has its headroom, and you can still keep a browser open. It assumes the memory is dedicated to the model or close to it: either a graphics card’s VRAM, or system RAM with the OS and apps accounted for separately.
| Model class | Examples of the size | Q4 file, roughly | Memory for the model to be comfortable |
|---|---|---|---|
| 1–4B | Llama 3.2 3B, Phi-class, Gemma small variants | 1–2.5 GB | 4 GB |
| 7–9B | Llama 3.1 8B, Qwen 7B, Mistral 7B | 4.5–5.5 GB | 8 GB |
| 12–14B | Qwen 14B, Mistral Nemo 12B | 7.5–9 GB | 12 GB |
| 24–32B | Mistral Small 24B, Gemma 27B, Qwen 32B | 14–20 GB | 24 GB |
| 70B | Llama 3.x 70B, Qwen 72B | 40–43 GB | 48 GB minimum, 64 GB to breathe |
| 100B+ dense | Mistral Large-class | 65 GB and up | 96–128 GB |
Then add the rest of the machine on top if the model is living in system RAM. That gives the numbers people usually ask about:
- 8 GB total system RAM: a 3B model runs well. A 7B at Q4 technically loads, but only with a short context and almost nothing else open. This is a demo tier.
- 16 GB: 7–8B models are comfortable. 12–14B at Q4 load if you close the browser and keep the context short, which is exactly where my first attempt went wrong.
- 32 GB: 14B is comfortable with long context. 24–32B at Q4 fits with care. This is the first tier where the useful mid-size models live without negotiation.
- 64 GB: 32B with generous context, or a 70B at Q4 with a modest window.
- 128 GB: 70B at higher quality quantization or with long context, and the first taste of the 100B+ class.
The mixture-of-experts wrinkle
A newer family of models breaks the neat relationship between parameter count and usefulness. Mixture-of-experts (MoE) models have a large total parameter count but only use a fraction of it for each token. Qwen’s 30B-A3B, for example, has about 30 billion parameters in total but activates roughly 3 billion per token. OpenAI’s open-weight gpt-oss-20b ships as a file of about 13 GB and is designed to run in 16 GB.
For the chart, the important point is simple: MoE models need memory for all of their parameters, not just the active ones. A 30B MoE at Q4 still needs roughly 18 GB for the weights. What you gain is speed, because each token only touches a small slice of that memory. What you do not gain is a smaller box. I have watched people load a “3B active” model into an 8 GB machine on the strength of that number and hit the same wall I hit with my 13B.

Where the chart lies a little
Smaller quants are not free memory
You can fit a bigger model in the same box by dropping to Q3 or Q2. The chart does not stop you. The model will, by getting noticeably worse at following instructions, keeping facts straight and producing clean code. In my experience a 14B at Q4 beats a 32B squeezed to Q2 for nearly every everyday job. Go below Q4 only when you have tested the result on your own prompts, not because a forum post said it loaded.
The KV cache can be quantized too
llama.cpp and tools built on it can store the context cache at 8-bit instead of 16-bit, which roughly halves that line of the bill. It is the cheapest long-context trick available and usually costs little quality. It does not change the weights line, so it will not rescue a model that is too big to begin with.
Unified memory changes who pays
On Apple Silicon there is no separate VRAM. The same pool serves macOS, your apps and the model, and the system limits how much of it the GPU can claim by default. So a 16 GB Mac is not a 16 GB graphics card. It is closer to the 16 GB laptop row above, with the OS taking its cut first. If you are sizing a Mac, the next thing to understand is why unified memory is the real constraint on Apple Silicon, because the ceiling there is not quite the number on the spec sheet.
“Fits” and “fast” are different questions
Everything in the chart is about whether the model can live in memory at all. It says nothing about how quickly it answers. A 32B model that fits in 32 GB of ordinary system RAM will run, but it may produce a few words a second, while the same model on a card with enough VRAM is several times faster, because generation speed is limited mostly by how fast memory can be read, not by how much of it there is. That trade-off is real, and it is the whole argument when you are deciding what to buy. It is also a separate argument from this one.
How I size a machine now
After the 13B episode I stopped starting with the model I wanted and started with the job:
- Name the job. Rewriting emails and summarizing notes is a 7–8B job. Decent coding help and longer reasoning is a 14–32B job. Anything approaching the big cloud models is 70B and up, and even then it will not match them.
- Take the chart row for that model class. Use the “comfortable” column, not the file size.
- Decide how much context you really use. If you paste whole documents or long code files, add the KV cost for 16K or 32K tokens rather than 8K.
- Add the rest of the machine if the model will share system RAM with your normal work. Be honest about the browser.
- Round up to the next real memory size. Nobody sells 27 GB. If the answer is 27, the answer is 32.
For most people who want a genuinely useful local assistant rather than a party trick, that process lands on 32 GB of system memory, or a graphics card with 12 to 16 GB of VRAM for the 8–14B class and 24 GB for the 32B class. 16 GB of system RAM is a perfectly good way to learn. 8 GB is a way to find out you are interested.
The short version
Take the parameter count, multiply by 0.6 for a Q4 model, add a gigabyte or two for context and runtime, then add whatever the rest of your computer is already using. If that total is more than the memory you have, the model does not fit, however confidently the download page states the file size. I learned that from a 7.9 GB file and a laptop that sounded like a hair dryer. The chart is cheaper.