The KV cache for Llama 3.1 8B takes 128 KiB per token at BF16 — exactly 1 GiB for an 8,192-token conversation, 16 GiB at the full 128K context. The size is layers × KV heads × head dimension × 2 × bytes per element, multiplied by tokens and by concurrent sequences. Every one of those numbers is in the model’s config.json.
Why the model keeps a cache at all
A transformer generates one token at a time, and to pick each one it attends over every token before it. Attention needs a key and a value vector for each of those earlier tokens, in every layer. Without a cache the model would recompute all of them at every step — the cost of generating a reply would grow with the square of its length.
So the model computes each token’s keys and values once, on the step that token arrives, and keeps them. That store is the KV cache. It is the reason the first token of a reply takes noticeably longer than the rest (the prefill step has to build the cache for the whole prompt), and the reason a long chat gets slower and eventually runs out of memory: the cache grows by a fixed amount for every token seen and never shrinks until the sequence ends.
The cache lives in GPU memory next to the weights, because every generated token reads all of it. The question this page answers is how much of that memory it takes.
The formula
Per token, per layer, a grouped-query or multi-head model stores one key and one value vector for each KV head:
elements per layer e = 2 × H_kv × D
bytes per token = L × e × bytes_per_element
per sequence = bytes per token × S
total = per sequence × B- L —
num_hidden_layers - H_kv —
num_key_value_heads. Notnum_attention_heads; see below. - D —
head_dim. If the config does not list it, it ishidden_size ÷ num_attention_heads. - 2 — one key and one value.
- bytes per element — 4 for FP32, 2 for FP16/BF16, 1 for FP8 or INT8, 0.5 for INT4 or FP4. This is the cache precision, which serving engines set separately from the weights (
--kv-cache-dtype fp8in vLLM). - S — context length: prompt plus generated tokens.
- B — concurrent sequences.
Nothing here is an estimate. The cache is a tensor of known shape, so unlike a whole-model VRAM figure there is no overhead multiplier — the allocator slack is a caveat, not a term.
Multi-head latent attention (DeepSeek V2 and V3) does not store keys and values at all. It stores one compressed latent vector per token per layer, plus a small decoupled key that carries the position encoding:
elements per layer e = d_c + d_Rwhere d_c is kv_lora_rank and d_R is qk_rope_head_dim. There is no factor of 2, and the head count does not appear. Attention is reconstructed from the latent at read time, which trades a little compute for a cache that is tens of times smaller.
A worked example, step by step
Llama 3.1 8B, from its config.json: 32 layers, 8 KV heads, head dimension 128. BF16 cache, 8,192 tokens, one sequence.
e = 2 × 8 × 128 = 2,048 elements per layer
bytes per token = 32 × 2,048 × 2 = 131,072 bytes = 128 KiB
per sequence = 131,072 × 8,192 = 1,073,741,824 = 1.00 GiBFour concurrent sequences at 4,096 tokens is the same 131,072 × 4,096 × 4 = 2,147,483,648 bytes = 2.00 GiB. Each sequence pays for its own context; halving the context and quadrupling the users doubles the cache.
Llama 3.1 70B keeps the same 8 KV heads and head dimension but has 80 layers, so it is 2.5× the 8B figure per token:
bytes per token = 80 × 2,048 × 2 = 327,680 bytes = 320 KiB
at 8,192 tokens = 2.50 GiB
at 131,072 tokens = 40.00 GiBThat 40 GiB is why a 70B model whose FP8 weights fit on one 80 GB card still cannot serve its full context to more than a handful of users. At FP8 the cache halves to 20 GiB; at INT4 it is 10 GiB.
DeepSeek V3, MLA: 61 layers, kv_lora_rank 512, qk_rope_head_dim 64. FP8 cache, 32,768 tokens.
e = 512 + 64 = 576 elements per layer
bytes per token = 61 × 576 × 1 = 35,136 bytes = 34.3 KiB
per sequence = 35,136 × 32,768 = 1,151,336,448 = 1.07 GiBWith its 128 attention heads at head dimension 128, the standard formula would give 32,768 elements per layer instead of 576 — about 57× more. That ratio is what MLA buys.
Why does everyone’s KV cache formula give a different answer?
Most blog posts still use the pre-2023 formula: 2 × layers × hidden_size × bytes. It was right when every model was multi-head attention, because hidden_size equals attention_heads × head_dim and every attention head had its own key and value.
Grouped-query attention broke it. Llama 3.1 8B has 32 attention heads but only 8 KV heads: several query heads share one key/value pair, and only the KV count decides the cache. hidden_size is 4,096; kv_heads × head_dim is 1,024. The old formula overstates the cache by exactly 4×, and on Llama 3.1 70B (64 attention heads, 8 KV heads) by 8×. If a page tells you Llama 3.1 8B needs 4 GiB of cache at 8K context, that is the mistake it made.
old: 2 × 32 × 4096 × 2 = 524,288 bytes per token (wrong for GQA)
new: 2 × 32 × 1024 × 2 = 131,072 bytes per tokenThe second most common discrepancy is units. 230 bytes is 1.00 GiB and 1.07 GB, and many pages mix the two in one table. This tool reports GiB throughout, because that is what nvidia-smi and your serving engine’s logs report.
Does the KV cache depend on parameter count?
No — and the counter-examples are worth knowing. The cache depends on layers, KV heads and head dimension, and models of the same size differ on all three:
| Model | Layers | KV heads | Head dim | Per token (BF16) | Per 1,024 tokens |
|---|---|---|---|---|---|
| Qwen 2.5 7B | 28 | 4 | 128 | 56 KiB | 56 MiB |
| Llama 3.1 8B | 32 | 8 | 128 | 128 KiB | 128 MiB |
| Mistral 7B | 32 | 8 | 128 | 128 KiB | 128 MiB |
| Phi-4 14B | 40 | 10 | 128 | 200 KiB | 200 MiB |
| Yi 34B | 60 | 8 | 128 | 240 KiB | 240 MiB |
| Llama 3.1 70B | 80 | 8 | 128 | 320 KiB | 320 MiB |
| Gemma 2 9B | 42 | 8 | 256 | 336 KiB | 336 MiB |
| Gemma 2 27B | 46 | 16 | 128 | 368 KiB | 368 MiB |
| DeepSeek V3 (MLA) | 61 | — | — | 68.6 KiB | 68.6 MiB |
Qwen 2.5 7B has four KV heads and 28 layers, so it caches less than half of what Llama 3.1 8B does per token — the same card serves twice the users. Gemma 2 9B goes the other way: its head dimension is 256, not the 224 you get by dividing hidden size by head count, so a 9B model carries a bigger cache than a 70B one. Mixture-of-experts models are not special here: Mixtral 8x7B has the same attention as Mistral 7B and the same 128 KiB per token, because the experts are in the feed-forward blocks, not the attention.
What is a good KV cache size?
It is what is left, not a target. Take the card, subtract the weights and the framework’s fixed overhead, and that remainder divided by bytes-per-token is your capacity — in tokens for one user, or in users at a given context. The budget field on this page does that division.
For Llama 3.1 8B on a 24 GB card: about 15 GiB of BF16 weights, roughly a gigabyte of CUDA context and buffers, so around 8 GiB for cache. That is eight users at 8K context, or one user at 64K. Switch the cache to FP8 and it is sixteen users, or 128K.
This is also what vLLM’s startup line means. It reserves a fixed share of the card (gpu_memory_utilization, default 0.9), subtracts what the weights need, and reports the rest as # GPU blocks: N — each block is 16 tokens at the per-token figure above. N × 16 is your total token capacity across all sequences. Prefix caching stores a shared system prompt once, so real capacity can be higher than this page’s figure when many requests start the same way.
What this does not include
- Sliding-window layers. Gemma 2 alternates local layers with a 4,096-token window and global ones; Mistral 7B v0.1 used a 4K window on every layer. A local layer’s cache stops growing at the window, so at long context the formula overstates those models — for Gemma 2 at 8K, by roughly a quarter. Enter the full layer count and treat the result as an upper bound.
- Allocator granularity. Engines allocate in fixed blocks (16 tokens in vLLM), so real usage rounds each sequence up to a block boundary.
- Prefix caching. Assumed off; every sequence pays for its full context.
- Speculative decoding, beam search, n-best sampling. Each candidate carries its own cache — enter the candidate count as sequences.
- Encoder–decoder models (Whisper, T5) keep a cross-attention cache of a different shape.
- Per-channel or grouped KV quantisation (KIVI and similar). Only whole-tensor precisions are modelled; treat the INT4 column as a floor.
- Offload. Spilling the cache to CPU RAM or NVMe (vLLM CPU offload, LMCache) changes where it lives, not how big it is — and decode slows to the speed of fetching it back.
- Weights, activations, CUDA context. None of it. For the whole-model figure and which GPUs fit, use the LLM VRAM calculator.