DeepLearn Tools
Menu
AI & LLM Engineering

LLM VRAM Calculator

Llama 3.1 8B at BF16 needs about 19 GiB with an 8K context — not the 16 GB its download suggests. See weights, KV cache and overhead separately, and which cards fit.

By Seifeur Guizeni, Sr. Agentic AI Architect & AI/ML Consultant

num_key_value_heads, not attention heads.

Weights
14.96 GiB
KV cache
1.00 GiB
Context + overhead
3.19 GiB

Total VRAM

19.15 GiB

Smallest card that fits: Tesla A10 — about 25 tok/s

Which GPUs fit

A card is usable to about 92% of its advertised capacity — the rest is display output, driver reserve and allocator fragmentation.

GPUVRAMFitsSpeed
RTX 3060 12GB Entry point for local models. 12 GiBNo—
RTX 2080 Ti No BF16. 11 GiBNo—
RTX 2080 Ti (22GB mod) Community hardware mod. No BF16. 22 GiBFits~25 tok/s
Intel Arc A770 16GB SYCL / IPEX-LLM only. 16 GiBNo—
RTX 4060 Ti 16GB Fits a lot, reads it slowly. 16 GiBNo—
RTX A4000 Single slot, 140W. 16 GiBNo—
RTX 4070 Ti Super 16 GiBNo—
RTX 4080 / Super 16 GiBNo—
RTX 5080 16 GiBNo—
RTX 4000 Ada 70W small form factor. 20 GiBNo—
RX 7900 XT ROCm. 20 GiBNo—
Tesla P40 No tensor cores. No tensor cores — FP16 runs very slowly on this card. INT8 is its usable path.24 GiBFits~14 tok/s
Tesla A10 24 GiBFits~25 tok/s
RTX A5000 NVLink. 24 GiBFits~31 tok/s
RTX 3090 / 3090 Ti NVLink. No FP8. 24 GiBFits~38 tok/s
RX 7900 XTX ROCm. 24 GiBFits~39 tok/s
RTX 4090 No NVLink. 24 GiBFits~41 tok/s
Tesla V100 32GB HBM2. No BF16, no INT4. 32 GiBFits~37 tok/s
RTX 5090 32 GiBFits~73 tok/s
RTX A6000 NVLink. 48 GiBFits~31 tok/s
RTX 6000 Ada Largest single workstation card. 48 GiBFits~39 tok/s
2× RTX 3090 48 GiBFits~38 tok/s
2× RTX 4090 48 GiBFits~41 tok/s
2× RTX A6000 NVLink pools the memory. 96 GiBFits~31 tok/s
4× RTX 3090 96 GiBFits~38 tok/s

Speed is a bandwidth ceiling for decoding at batch 1, not a benchmark. It ignores prompt processing and overstates cards without tensor cores.

Share this tool

Llama 3.1 8B at BF16 with an 8,192-token context needs about 19.15 GiB of VRAM — so a 24 GB card, not the 16 GB one its 16.07 GB download suggests. Quantise the same model to 4-bit and it drops to 5.69 GiB, which a 12 GB RTX 3060 handles comfortably. The gap between those two numbers is what this calculator exists to show.

The formula

Everything is computed in bytes and converted once for display. Card capacities are binary — a “24 GB” 4090 is 24 GiB, or 25.77 billion bytes — while model files on Hugging Face are quoted in decimal GB. Mixing the two is the most common way a VRAM estimate goes wrong, and it is wrong by about 7%: enough to promise a fit that does not happen.

bytes_per_param = 4 (FP32) | 2 (FP16/BF16) | 1 (FP8, INT8) | 0.5 (INT4)

weights   = params × bytes_per_param
kv_cache  = 2 × layers × kv_heads × head_dim × context × batch × kv_bytes
total     = (weights + kv_cache) × 1.2

The leading 2 in the KV term is K and V. The 1.2 covers the CUDA context, activation buffers and allocator fragmentation; it is the only estimate in the inference path, and it is fitted to typical PyTorch behaviour rather than derived.

Where it breaks down. A context length of zero, or a batch size of zero, returns an error rather than a weights-only figure. Both would produce a clean and plausible number — the weights plus overhead — for a configuration that cannot generate a single token, and a calculator that answers an impossible question confidently is worse than one that refuses.

Fine-tuning

Training holds three things the weights alone do not: gradients, optimizer state, and activations kept for the backward pass.

Component Full LoRA QLoRA
Base weights 2 B/param 2 B/param 0.5 B/param
Gradients 2 B/param 2 × trainable 2 × trainable
Optimizer (AdamW: fp32 momentum + variance) 8 B/param 8 × trainable 8 × trainable
base        = params × base_bytes
trainable   = params × trainable_fraction     (= params for full fine-tuning)
grad_opt    = trainable × 10
activations = layers × seq × batch × hidden × 2 × k
              k = 1 with gradient checkpointing, 12 without
total       = base + grad_opt + activations

A LoRA that trains 100% of the parameters is full fine-tuning, and the formula converges there on its own — there is no special case, which is the honest result rather than a coincidence.

A worked example, step by step

Llama 3.1 8B, BF16, 8,192 tokens, batch 1. The dimensions come from that model’s config.json: 32 layers, 8 KV heads, head dimension 128.

weights  = 8.03e9 × 2                        = 16.06e9 bytes = 14.96 GiB
kv_cache = 2 × 32 × 8 × 128 × 8192 × 1 × 2   =  1.07e9 bytes =  1.00 GiB
total    = (14.96 + 1.00) × 1.2                              = 19.15 GiB

That does not fit a 16 GiB RTX 4060 Ti. It fits a 24 GiB RTX 3090 with room to spare. Drop the weights to INT4 while keeping the cache at FP16, as llama.cpp does by default, and the weights fall to 3.74 GiB and the total to 5.69 GiB — a 12 GiB 3060 handles it.

The same model with QLoRA, sequence length 2,048, batch 1, 0.5% trainable, checkpointing on:

base        = 8.03e9 × 0.5                   = 3.74 GiB
trainable   = 8.03e9 × 0.005                 = 40.2M parameters
grad_opt    = 40.2e6 × 10                    = 0.37 GiB
activations = 32 × 2048 × 1 × 4096 × 2 × 1   = 0.50 GiB
total                                        = 4.61 GiB

Turn gradient checkpointing off and activations jump to 6.0 GiB, taking the total to 10.1 GiB. That is why every “QLoRA on a 12 GB card” guide has “enable gradient checkpointing” attached to it.

Why does my GPU need more VRAM than the model file size?

This is the question most people are actually asking, and most pages answer it with a compatibility table instead of a reason. There are three components, and the download is only the first.

The weights are the floor. A 16.07 GB BF16 download is 16.07 GB of VRAM before anything else happens. That part is predictable, and it is the part every “model size” table shows.

The KV cache grows with the conversation. Every token the model has seen leaves a key and a value vector in memory, in every layer, and they stay there for the rest of the session. At 8K context on Llama 3.1 8B that is 1.00 GiB — a rounding error next to the weights. At 128K context on the same model it is 16 GiB, larger than the weights themselves. The cache is why a model that loaded fine crashes forty minutes into a long chat.

The framework holds back a fixed chunk. CUDA context, activation buffers, and memory the allocator has fragmented past usefulness. Roughly 20% on top of weights plus cache, which is the 1.2 in the formula.

Put together: 16 GB of weights wants about 20 GB of card. The rule of thumb worth remembering is that the download tells you the floor, never the requirement.

Why does this ask for KV heads instead of attention heads?

There is a formula circulating in blog posts, Stack Overflow answers and several other calculators that uses hidden_size where the one above uses kv_heads × head_dim:

kv_cache = 2 × layers × hidden_size × context × batch × bytes    ← the old form

That was correct for multi-head attention, where every query head carried its own key and value. It is wrong for every model shipped with grouped-query attention since Llama 2 70B, and GQA is now close to universal.

Llama 3.1 8B has 32 attention heads and a hidden size of 4,096 — but only 8 KV heads. Eight heads at dimension 128 is 1,024, not 4,096. So for the example above:

Formula KV cache at 8K context
hidden_size (the old form) 4.00 GiB
kv_heads × head_dim (GQA) 1.00 GiB

A factor of four, in the direction of telling you to buy a bigger card than you need. If you check this tool against a page using the old formula, that is the discrepancy you will find, and this is the side that matches reality: the working figure in the field — roughly 0.5 to 2 GB per session at 8K context — is the GQA number. The hidden-size version does not match anyone’s observed usage, which is the tell.

Grouped-query attention shares one key/value pair across a group of query heads. It exists precisely to shrink this cache, and a calculator that ignores it is measuring a model nobody ships any more.

VRAM decides what loads. Bandwidth decides whether it is usable.

Most “best GPU for LLMs” advice stops at capacity, and capacity is only half the answer.

Decoding is memory-bound. To generate one token, the GPU reads every weight and the entire KV cache — then does it again for the next token. Nothing about that is compute-limited on a modern card, so the ceiling is simply bandwidth divided by bytes read:

tokens_per_sec ≈ (bandwidth_GB/s × 0.7) / (weights + kv_cache in GB)

The 0.7 is measured efficiency, not physics, and it is the loosest number in this tool — which is why speeds here carry no decimal place.

The RTX 4060 Ti and the RTX 4080 both hold 16 GB. They fit exactly the same models. The 4060 Ti reads its memory at 288 GB/s and the 4080 at 717 GB/s, so the 4080 generates roughly 2.5× faster on identical work. Two cards, same fit column, completely different experience.

This also explains the thing every local-LLM user notices and few pages state plainly: generation gets slower as the conversation gets longer. The KV cache is in the denominator. At 8K context on an 8B model it adds a few percent to the bytes read per token. At 128K it can be most of what the card is reading, and the same model on the same card generates at a fraction of its opening speed — without ever running out of memory.

Why 4-bit is not free

Quantisation is the single biggest lever here — 4× off the weights, moving a 70B model from “rent a datacenter GPU” to “two used 3090s”. It is not lossless, and the loss is not evenly distributed.

Some models tolerate it well. Llama models hold up at 4-bit for most tasks. Gemma 2 degrades noticeably at the same bit width, and models trained with a lot of knowledge packed into few parameters tend to suffer most. Reasoning and code generation degrade before conversational fluency does, so a quantised model can sound completely fine while getting measurably worse at the thing you wanted it for.

Formats are not interchangeable either:

Format Best for Notes
GGUF Apple Silicon, llama.cpp, CPU offload The only practical option on a Mac
AWQ / FP8 vLLM and SGLang servers Fast batched serving on NVIDIA
EXL2 Consumer NVIDIA, single user Flexible bit rates, very fast
NF4 (bitsandbytes) QLoRA fine-tuning What the QLoRA mode above assumes

One number worth carrying: GGUF quants are not clean bit widths. Q4_K_M averages about 4.8 bits per weight, not 4.0, because it keeps attention and embedding tensors at higher precision. The INT4 option in this calculator models a uniform 4-bit format, so it will underestimate a Q4_K_M file by 15–20%. Add that back before deciding a card is big enough.

Do MoE models only need memory for the active experts?

No, and this catches people out constantly. Mixtral 8x7B has 46.7 billion parameters. Roughly 13 billion of them are used for any given token — but all 46.7 billion sit in VRAM, because the router picks different experts for the next token and there is no time to swap them in.

Mixture-of-experts buys you compute, not memory. Enter the total parameter count, not the active one. The saving is real, it just shows up in speed rather than in capacity.

What this does not include

The 1.2 inference multiplier and the activation constant k are estimates fitted to typical PyTorch behaviour. Every other term in both formulas is exact arithmetic on figures from the model’s own configuration.

Frequently asked questions

Why does my GPU need more VRAM than the model file size?
The weights are the floor, not the requirement. The KV cache grows with every token in the conversation, and the framework holds back a fixed chunk for CUDA context and activation buffers. As a rule, 16 GB of weights wants about 20 GB of card.
How much VRAM do I need to run Llama 3.1 8B?
About 19.15 GiB at BF16 with an 8K context, so a 24 GB card. Quantised to 4-bit it drops to 5.69 GiB and a 12 GB RTX 3060 is enough. The download is 16.07 GB, which is why the 16 GB cards look like they should work and do not.
Why does the calculator ask for KV heads instead of attention heads?
Grouped-query attention shares one key/value pair across several query heads, and only the KV count drives cache size. Llama 3.1 8B has 32 attention heads but 8 KV heads. The older formula that uses hidden size instead overstates its cache by exactly 4×.
Does context length matter more than parameter count?
Not until it does. At 8K context the KV cache is a rounding error next to the weights — 1 GiB against 15. At 128K on a 70B model the cache is larger than the weights themselves, and it is what decides whether the session survives.
Can I fine-tune a 7B model on a 16 GB card?
With QLoRA and gradient checkpointing, yes — around 4.6 GiB at sequence length 2048. Turn checkpointing off and the same job needs 10.1 GiB. Full fine-tuning the same model needs roughly 100 GB, which is datacenter hardware.
Why is my model slower in a long conversation than a short one?
The GPU re-reads the whole KV cache for every token it generates, so the cache is in the denominator of the speed estimate. At 8K context on an 8B model that costs a few percent. At 128K it can be most of what the card is reading.
Do MoE models like Mixtral only need memory for the active experts?
No. All 46.7 billion parameters of Mixtral 8x7B sit in VRAM even though only about 13 billion run per token, because the router picks different experts for the next token. Mixture-of-experts buys compute, not memory. Enter the total count.
Why does this not match what nvidia-smi shows under vLLM?
vLLM and SGLang claim a fixed share of the card at startup — gpu_memory_utilization, default 0.9 — and fill the rest with KV cache blocks. What you are reading is their reservation, not the model’s need, so the two figures will never agree.