DeepLearn Tools
Menu
AI & LLM Engineering

KV Cache Calculator

Llama 3.1 8B caches 128 KiB per token — exactly 1 GiB at 8K context, 16 GiB at 128K. Enter the model’s config.json numbers and see the cache per token, per sequence and in total, plus how many sequences fit your memory.

By Seifeur Guizeni, Sr. Agentic AI Architect & AI/ML Consultant

num_hidden_layers

num_key_value_heads, not attention heads. For plain MHA the two are equal.

head_dim, or hidden size ÷ attention heads if absent.

Prompt plus generated tokens.

What is left after the weights and overhead — not the card size.

Per token
128.0 KiB 131,072 bytes
Per sequence
1.00 GiB

KV cache total

1.00 GiB

Same total at each precision

FP16 / BF16
1.00 GiB
FP8 / INT8
512.00 MiB
INT4 / FP4
256.00 MiB

Within 8.00 GiB

8 concurrent sequences at 8,192 tokens, or a single sequence of up to 65,536 tokens.

Binary units throughout: 1 GiB = 230 bytes, what nvidia-smi reports. In decimal GB the figures are about 7% larger.

Share this tool

The KV cache for Llama 3.1 8B takes 128 KiB per token at BF16 — exactly 1 GiB for an 8,192-token conversation, 16 GiB at the full 128K context. The size is layers × KV heads × head dimension × 2 × bytes per element, multiplied by tokens and by concurrent sequences. Every one of those numbers is in the model’s config.json.

Why the model keeps a cache at all

A transformer generates one token at a time, and to pick each one it attends over every token before it. Attention needs a key and a value vector for each of those earlier tokens, in every layer. Without a cache the model would recompute all of them at every step — the cost of generating a reply would grow with the square of its length.

So the model computes each token’s keys and values once, on the step that token arrives, and keeps them. That store is the KV cache. It is the reason the first token of a reply takes noticeably longer than the rest (the prefill step has to build the cache for the whole prompt), and the reason a long chat gets slower and eventually runs out of memory: the cache grows by a fixed amount for every token seen and never shrinks until the sequence ends.

The cache lives in GPU memory next to the weights, because every generated token reads all of it. The question this page answers is how much of that memory it takes.

The formula

Per token, per layer, a grouped-query or multi-head model stores one key and one value vector for each KV head:

elements per layer   e = 2 × H_kv × D
bytes per token      = L × e × bytes_per_element
per sequence         = bytes per token × S
total                = per sequence × B

Nothing here is an estimate. The cache is a tensor of known shape, so unlike a whole-model VRAM figure there is no overhead multiplier — the allocator slack is a caveat, not a term.

Multi-head latent attention (DeepSeek V2 and V3) does not store keys and values at all. It stores one compressed latent vector per token per layer, plus a small decoupled key that carries the position encoding:

elements per layer   e = d_c + d_R

where d_c is kv_lora_rank and d_R is qk_rope_head_dim. There is no factor of 2, and the head count does not appear. Attention is reconstructed from the latent at read time, which trades a little compute for a cache that is tens of times smaller.

A worked example, step by step

Llama 3.1 8B, from its config.json: 32 layers, 8 KV heads, head dimension 128. BF16 cache, 8,192 tokens, one sequence.

e               = 2 × 8 × 128        = 2,048 elements per layer
bytes per token = 32 × 2,048 × 2     = 131,072 bytes  = 128 KiB
per sequence    = 131,072 × 8,192    = 1,073,741,824  = 1.00 GiB

Four concurrent sequences at 4,096 tokens is the same 131,072 × 4,096 × 4 = 2,147,483,648 bytes = 2.00 GiB. Each sequence pays for its own context; halving the context and quadrupling the users doubles the cache.

Llama 3.1 70B keeps the same 8 KV heads and head dimension but has 80 layers, so it is 2.5× the 8B figure per token:

bytes per token = 80 × 2,048 × 2     = 327,680 bytes  = 320 KiB
at 8,192 tokens                      = 2.50 GiB
at 131,072 tokens                    = 40.00 GiB

That 40 GiB is why a 70B model whose FP8 weights fit on one 80 GB card still cannot serve its full context to more than a handful of users. At FP8 the cache halves to 20 GiB; at INT4 it is 10 GiB.

DeepSeek V3, MLA: 61 layers, kv_lora_rank 512, qk_rope_head_dim 64. FP8 cache, 32,768 tokens.

e               = 512 + 64           = 576 elements per layer
bytes per token = 61 × 576 × 1       = 35,136 bytes   = 34.3 KiB
per sequence    = 35,136 × 32,768    = 1,151,336,448  = 1.07 GiB

With its 128 attention heads at head dimension 128, the standard formula would give 32,768 elements per layer instead of 576 — about 57× more. That ratio is what MLA buys.

Why does everyone’s KV cache formula give a different answer?

Most blog posts still use the pre-2023 formula: 2 × layers × hidden_size × bytes. It was right when every model was multi-head attention, because hidden_size equals attention_heads × head_dim and every attention head had its own key and value.

Grouped-query attention broke it. Llama 3.1 8B has 32 attention heads but only 8 KV heads: several query heads share one key/value pair, and only the KV count decides the cache. hidden_size is 4,096; kv_heads × head_dim is 1,024. The old formula overstates the cache by exactly 4×, and on Llama 3.1 70B (64 attention heads, 8 KV heads) by 8×. If a page tells you Llama 3.1 8B needs 4 GiB of cache at 8K context, that is the mistake it made.

old:  2 × 32 × 4096 × 2 = 524,288 bytes per token   (wrong for GQA)
new:  2 × 32 × 1024 × 2 = 131,072 bytes per token

The second most common discrepancy is units. 230 bytes is 1.00 GiB and 1.07 GB, and many pages mix the two in one table. This tool reports GiB throughout, because that is what nvidia-smi and your serving engine’s logs report.

Does the KV cache depend on parameter count?

No — and the counter-examples are worth knowing. The cache depends on layers, KV heads and head dimension, and models of the same size differ on all three:

Model Layers KV heads Head dim Per token (BF16) Per 1,024 tokens
Qwen 2.5 7B 28 4 128 56 KiB 56 MiB
Llama 3.1 8B 32 8 128 128 KiB 128 MiB
Mistral 7B 32 8 128 128 KiB 128 MiB
Phi-4 14B 40 10 128 200 KiB 200 MiB
Yi 34B 60 8 128 240 KiB 240 MiB
Llama 3.1 70B 80 8 128 320 KiB 320 MiB
Gemma 2 9B 42 8 256 336 KiB 336 MiB
Gemma 2 27B 46 16 128 368 KiB 368 MiB
DeepSeek V3 (MLA) 61 — — 68.6 KiB 68.6 MiB

Qwen 2.5 7B has four KV heads and 28 layers, so it caches less than half of what Llama 3.1 8B does per token — the same card serves twice the users. Gemma 2 9B goes the other way: its head dimension is 256, not the 224 you get by dividing hidden size by head count, so a 9B model carries a bigger cache than a 70B one. Mixture-of-experts models are not special here: Mixtral 8x7B has the same attention as Mistral 7B and the same 128 KiB per token, because the experts are in the feed-forward blocks, not the attention.

What is a good KV cache size?

It is what is left, not a target. Take the card, subtract the weights and the framework’s fixed overhead, and that remainder divided by bytes-per-token is your capacity — in tokens for one user, or in users at a given context. The budget field on this page does that division.

For Llama 3.1 8B on a 24 GB card: about 15 GiB of BF16 weights, roughly a gigabyte of CUDA context and buffers, so around 8 GiB for cache. That is eight users at 8K context, or one user at 64K. Switch the cache to FP8 and it is sixteen users, or 128K.

This is also what vLLM’s startup line means. It reserves a fixed share of the card (gpu_memory_utilization, default 0.9), subtracts what the weights need, and reports the rest as # GPU blocks: N — each block is 16 tokens at the per-token figure above. N × 16 is your total token capacity across all sequences. Prefix caching stores a shared system prompt once, so real capacity can be higher than this page’s figure when many requests start the same way.

What this does not include

Frequently asked questions

How do I calculate KV cache size?
Layers × KV heads × head dimension × 2 (keys and values) × bytes per element gives the cache per token; multiply by context length and by concurrent sequences. Llama 3.1 8B at BF16 is 128 KiB per token, so 8K tokens is exactly 1 GiB. All four architecture numbers are in the model’s config.json.
Which numbers do I take from config.json?
num_hidden_layers, num_key_value_heads and head_dim. Not num_attention_heads — on a grouped-query model that is 4–8× larger than the KV count and gives a 4–8× wrong answer. If head_dim is missing, it is hidden_size ÷ num_attention_heads.
How big is the KV cache for Llama 3 70B?
320 KiB per token at BF16: 2.5 GiB for an 8K conversation, 40 GiB at the full 128K context. At FP8 halve both. That 40 GiB is why a 70B model whose weights fit two 80 GB cards still cannot serve its full context to many users at once.
Is the KV cache stored on the GPU?
Yes. It has to be read in full for every generated token, so it lives in VRAM next to the weights. Engines can spill it to CPU RAM or NVMe (vLLM CPU offload, LMCache), but the spilled part comes back over PCIe and decoding slows down accordingly.
Does the KV cache scale with parameter count?
No. It depends only on layers, KV heads and head dimension. Qwen 2.5 7B has 4 KV heads and 28 layers, so its cache is 56 KiB per token — less than half of Llama 3.1 8B’s 128 KiB for a model of the same size. Two 7B models can differ 2× in how many users they serve.
What is a good KV cache size?
It is what is left, not a target. Take the card, subtract the weights and the framework’s fixed overhead, and divide the remainder by bytes per token per sequence — that is your concurrency at a given context. The memory budget field on this page does that division.
Does quantising the KV cache to FP8 hurt quality?
Barely, in most published evaluations — it is the cheapest halving available and every major serving engine supports it. INT4 KV is noticeably lossier and usually needs a per-channel scheme; treat the INT4 column as a floor, not a recommendation.
Is this the same thing as L1, L2 or L3 cache?
No. Those are small SRAM caches inside a CPU or GPU die, measured in kilobytes to megabytes, that hold recently used data. The KV cache is a tensor the model itself builds in GPU memory, measured in gigabytes. Same word, unrelated thing.