Llama 3.1 8B at BF16 with an 8,192-token context needs about 19.15 GiB of VRAM — so a 24 GB card, not the 16 GB one its 16.07 GB download suggests. Quantise the same model to 4-bit and it drops to 5.69 GiB, which a 12 GB RTX 3060 handles comfortably. The gap between those two numbers is what this calculator exists to show.
The formula
Everything is computed in bytes and converted once for display. Card capacities are binary — a “24 GB” 4090 is 24 GiB, or 25.77 billion bytes — while model files on Hugging Face are quoted in decimal GB. Mixing the two is the most common way a VRAM estimate goes wrong, and it is wrong by about 7%: enough to promise a fit that does not happen.
bytes_per_param = 4 (FP32) | 2 (FP16/BF16) | 1 (FP8, INT8) | 0.5 (INT4)
weights = params × bytes_per_param
kv_cache = 2 × layers × kv_heads × head_dim × context × batch × kv_bytes
total = (weights + kv_cache) × 1.2- params — total parameter count, not the active count. See the note on MoE models below.
- layers —
num_hidden_layersfrom the model’sconfig.json. - kv_heads —
num_key_value_heads. Notnum_attention_heads, and the difference is large. - head_dim —
head_dimwhere the config declares one, otherwise hidden size ÷ attention heads. Gemma 2 9B declares 256 where the division gives 224, so read it rather than derive it. - context — tokens held in the conversation, prompt and generation together.
- kv_bytes — the cache has its own precision. It does not have to match the weights.
The leading 2 in the KV term is K and V. The 1.2 covers the CUDA context, activation buffers and allocator fragmentation; it is the only estimate in the inference path, and it is fitted to typical PyTorch behaviour rather than derived.
Where it breaks down. A context length of zero, or a batch size of zero, returns an error rather than a weights-only figure. Both would produce a clean and plausible number — the weights plus overhead — for a configuration that cannot generate a single token, and a calculator that answers an impossible question confidently is worse than one that refuses.
Fine-tuning
Training holds three things the weights alone do not: gradients, optimizer state, and activations kept for the backward pass.
| Component | Full | LoRA | QLoRA |
|---|---|---|---|
| Base weights | 2 B/param | 2 B/param | 0.5 B/param |
| Gradients | 2 B/param | 2 × trainable | 2 × trainable |
| Optimizer (AdamW: fp32 momentum + variance) | 8 B/param | 8 × trainable | 8 × trainable |
base = params × base_bytes
trainable = params × trainable_fraction (= params for full fine-tuning)
grad_opt = trainable × 10
activations = layers × seq × batch × hidden × 2 × k
k = 1 with gradient checkpointing, 12 without
total = base + grad_opt + activationsA LoRA that trains 100% of the parameters is full fine-tuning, and the formula converges there on its own — there is no special case, which is the honest result rather than a coincidence.
A worked example, step by step
Llama 3.1 8B, BF16, 8,192 tokens, batch 1. The dimensions come from that model’s config.json: 32 layers, 8 KV heads, head dimension 128.
weights = 8.03e9 × 2 = 16.06e9 bytes = 14.96 GiB
kv_cache = 2 × 32 × 8 × 128 × 8192 × 1 × 2 = 1.07e9 bytes = 1.00 GiB
total = (14.96 + 1.00) × 1.2 = 19.15 GiBThat does not fit a 16 GiB RTX 4060 Ti. It fits a 24 GiB RTX 3090 with room to spare. Drop the weights to INT4 while keeping the cache at FP16, as llama.cpp does by default, and the weights fall to 3.74 GiB and the total to 5.69 GiB — a 12 GiB 3060 handles it.
The same model with QLoRA, sequence length 2,048, batch 1, 0.5% trainable, checkpointing on:
base = 8.03e9 × 0.5 = 3.74 GiB
trainable = 8.03e9 × 0.005 = 40.2M parameters
grad_opt = 40.2e6 × 10 = 0.37 GiB
activations = 32 × 2048 × 1 × 4096 × 2 × 1 = 0.50 GiB
total = 4.61 GiBTurn gradient checkpointing off and activations jump to 6.0 GiB, taking the total to 10.1 GiB. That is why every “QLoRA on a 12 GB card” guide has “enable gradient checkpointing” attached to it.
Why does my GPU need more VRAM than the model file size?
This is the question most people are actually asking, and most pages answer it with a compatibility table instead of a reason. There are three components, and the download is only the first.
The weights are the floor. A 16.07 GB BF16 download is 16.07 GB of VRAM before anything else happens. That part is predictable, and it is the part every “model size” table shows.
The KV cache grows with the conversation. Every token the model has seen leaves a key and a value vector in memory, in every layer, and they stay there for the rest of the session. At 8K context on Llama 3.1 8B that is 1.00 GiB — a rounding error next to the weights. At 128K context on the same model it is 16 GiB, larger than the weights themselves. The cache is why a model that loaded fine crashes forty minutes into a long chat.
The framework holds back a fixed chunk. CUDA context, activation buffers, and memory the allocator has fragmented past usefulness. Roughly 20% on top of weights plus cache, which is the 1.2 in the formula.
Put together: 16 GB of weights wants about 20 GB of card. The rule of thumb worth remembering is that the download tells you the floor, never the requirement.
Why does this ask for KV heads instead of attention heads?
There is a formula circulating in blog posts, Stack Overflow answers and several other calculators that uses hidden_size where the one above uses kv_heads × head_dim:
kv_cache = 2 × layers × hidden_size × context × batch × bytes ← the old formThat was correct for multi-head attention, where every query head carried its own key and value. It is wrong for every model shipped with grouped-query attention since Llama 2 70B, and GQA is now close to universal.
Llama 3.1 8B has 32 attention heads and a hidden size of 4,096 — but only 8 KV heads. Eight heads at dimension 128 is 1,024, not 4,096. So for the example above:
| Formula | KV cache at 8K context |
|---|---|
hidden_size (the old form) |
4.00 GiB |
kv_heads × head_dim (GQA) |
1.00 GiB |
A factor of four, in the direction of telling you to buy a bigger card than you need. If you check this tool against a page using the old formula, that is the discrepancy you will find, and this is the side that matches reality: the working figure in the field — roughly 0.5 to 2 GB per session at 8K context — is the GQA number. The hidden-size version does not match anyone’s observed usage, which is the tell.
Grouped-query attention shares one key/value pair across a group of query heads. It exists precisely to shrink this cache, and a calculator that ignores it is measuring a model nobody ships any more.
VRAM decides what loads. Bandwidth decides whether it is usable.
Most “best GPU for LLMs” advice stops at capacity, and capacity is only half the answer.
Decoding is memory-bound. To generate one token, the GPU reads every weight and the entire KV cache — then does it again for the next token. Nothing about that is compute-limited on a modern card, so the ceiling is simply bandwidth divided by bytes read:
tokens_per_sec ≈ (bandwidth_GB/s × 0.7) / (weights + kv_cache in GB)The 0.7 is measured efficiency, not physics, and it is the loosest number in this tool — which is why speeds here carry no decimal place.
The RTX 4060 Ti and the RTX 4080 both hold 16 GB. They fit exactly the same models. The 4060 Ti reads its memory at 288 GB/s and the 4080 at 717 GB/s, so the 4080 generates roughly 2.5× faster on identical work. Two cards, same fit column, completely different experience.
This also explains the thing every local-LLM user notices and few pages state plainly: generation gets slower as the conversation gets longer. The KV cache is in the denominator. At 8K context on an 8B model it adds a few percent to the bytes read per token. At 128K it can be most of what the card is reading, and the same model on the same card generates at a fraction of its opening speed — without ever running out of memory.
Why 4-bit is not free
Quantisation is the single biggest lever here — 4× off the weights, moving a 70B model from “rent a datacenter GPU” to “two used 3090s”. It is not lossless, and the loss is not evenly distributed.
Some models tolerate it well. Llama models hold up at 4-bit for most tasks. Gemma 2 degrades noticeably at the same bit width, and models trained with a lot of knowledge packed into few parameters tend to suffer most. Reasoning and code generation degrade before conversational fluency does, so a quantised model can sound completely fine while getting measurably worse at the thing you wanted it for.
Formats are not interchangeable either:
| Format | Best for | Notes |
|---|---|---|
| GGUF | Apple Silicon, llama.cpp, CPU offload | The only practical option on a Mac |
| AWQ / FP8 | vLLM and SGLang servers | Fast batched serving on NVIDIA |
| EXL2 | Consumer NVIDIA, single user | Flexible bit rates, very fast |
| NF4 (bitsandbytes) | QLoRA fine-tuning | What the QLoRA mode above assumes |
One number worth carrying: GGUF quants are not clean bit widths. Q4_K_M averages about 4.8 bits per weight, not 4.0, because it keeps attention and embedding tensors at higher precision. The INT4 option in this calculator models a uniform 4-bit format, so it will underestimate a Q4_K_M file by 15–20%. Add that back before deciding a card is big enough.
Do MoE models only need memory for the active experts?
No, and this catches people out constantly. Mixtral 8x7B has 46.7 billion parameters. Roughly 13 billion of them are used for any given token — but all 46.7 billion sit in VRAM, because the router picks different experts for the next token and there is no time to swap them in.
Mixture-of-experts buys you compute, not memory. Enter the total parameter count, not the active one. The saving is real, it just shows up in speed rather than in capacity.
What this does not include
- CPU and disk offload. A model that “doesn’t fit” here often still runs, slowly, with layers spilled to system RAM. llama.cpp does this by default.
- vLLM and SGLang reported usage. Both preallocate a fixed share of the card at startup (
gpu_memory_utilization, default 0.9) and fill it with KV blocks.nvidia-smiwill show that reservation, not the model’s need, so it will never match this figure. That is not a bug in either tool. - A headless machine’s extra headroom. The 92% usable figure assumes a display attached. On a headless Linux box with no desktop you get closer to 97%, which is why a 4090 sometimes runs a model this calculator says will not fit.
- Quantised KV cache schemes beyond FP8 and INT8. No per-channel or grouped variants.
- Multi-GPU overhead. The multi-card rows assume tensor parallelism and ignore activation replication and communication buffers — treat them as a floor. Layer-split offload, which is llama.cpp’s default, has a different and worse profile.
- Sharded training. DeepSpeed ZeRO, FSDP and multi-node setups change the fine-tuning arithmetic completely. The figures above are single-device.
- Laptop variants. Card capacities are the desktop figures. Mobile versions of the same SKU often ship less VRAM and far less bandwidth.
- Speed as a benchmark. The tokens/sec figure is a bandwidth ceiling with a fudge factor. It assumes batch 1, ignores prompt processing entirely, and will overstate cards without tensor cores — a Tesla P40 will miss it badly. Read it as “which of these is faster”, not as a measurement.
The 1.2 inference multiplier and the activation constant k are estimates fitted to typical PyTorch behaviour. Every other term in both formulas is exact arithmetic on figures from the model’s own configuration.