A transformer’s parameter count is the sum of its weight matrices: token embeddings (V × d), each layer’s attention and feed-forward network, the norms, and an output head unless it is tied to the embeddings. GPT-2 Small works out to exactly 124,439,808 parameters; Llama 3.1 8B to 8,030,261,248. The shortcut 12 · L · d² covers the layers alone.
The formula
Take the numbers from the model’s config.json:
- V —
vocab_size - d —
hidden_size(n_embdon GPT-2) - L —
num_hidden_layers - H, H_kv —
num_attention_headsandnum_key_value_heads(equal on plain multi-head attention) - h —
head_dim, usually d ÷ H - f —
intermediate_size, the feed-forward width (GPT-2 does not list it: it is 4d)
One transformer block holds:
attention = d·(H·h) + 2·d·(H_kv·h) + (H·h)·d Q, K and V, O
FFN = 2·d·f (standard: up, down)
= 3·d·f (gated / SwiGLU: gate, up, down)
norms = 2 × 2d (LayerNorm: a scale and a shift)
= 2 × d (RMSNorm: a scale only)plus biases where the architecture has them: H·h + 2·H_kv·h on Q, K and V, d on the attention output, and f (or 2f) + d on the FFN.
The whole model is:
total = V·d token embedding
+ P·d learned position table (0 for RoPE or ALiBi)
+ L × per-block all layers
+ final norm d or 2d
+ d·V LM head (0 when tied)The case people get wrong is the tied head. Most small models, GPT-2 and Qwen 2.5 0.5B included, reuse the embedding matrix as the output layer (tie_word_embeddings: true). It is one matrix used twice, so it is counted once. Adding d·V again overstates GPT-2 Small by 38.6 million — about 31%.
Where 12 · L · d² comes from
With plain multi-head attention, Q, K, V and O are each d × d: that is 4d². A standard FFN with f = 4d has two d × 4d matrices: 8d². Together that is 12d² per layer, and Kaplan et al.’s scaling-laws paper (2020) uses N ≈ 12 · L · d² for the non-embedding parameter count.
It is exact for GPT-2-style blocks, apart from biases and norms, and it drifts on modern ones:
- Grouped-query attention shrinks K and V. Llama 3.1 8B has 8 KV heads for 32 query heads, so its attention is 2.5d², not 4d².
- Gated FFNs use three matrices at an f that is rarely 4d. Llama 3.1 8B’s 3 × 4096 × 14,336 works out to 10.5d².
For Llama 3.1 8B the shortcut gives 6.44B for the layers; the exact non-embedding count is 7.50B. That is 14% short. For GPT-2 Small it is 0.1% short. The calculator shows both side by side.
A worked example, step by step
GPT-2 Small: V = 50,257, d = 768, L = 12, H = 12, f = 3,072, 1,024 learned positions, LayerNorm, biases on every linear layer, tied head.
token embedding 50,257 × 768 = 38,597,376
position table 1,024 × 768 = 786,432
per layer
attention 4 × 768² + 4 × 768 = 2,362,368
FFN 2 × 768 × 3,072 + 3,072 + 768 = 4,722,432
norms 2 × 2 × 768 = 3,072
----------
7,087,872
all layers 12 × 7,087,872 = 85,054,464
final norm 2 × 768 = 1,536
LM head tied = 0
----------
total = 124,439,808That is the “124M” on the model card. Pick the GPT-2 Small preset above and every row matches.
Llama 3.1 8B does the same arithmetic with GQA, a gated FFN, RMSNorm, no biases and an untied head:
token embedding 128,256 × 4,096 = 525,336,576
per layer
attention 2 × 4,096² + 2 × 4,096 × 1,024 = 41,943,040
FFN 3 × 4,096 × 14,336 = 176,160,768
norms 2 × 4,096 = 8,192
all layers 32 × 218,112,000 = 6,979,584,000
final norm = 4,096
LM head 4,096 × 128,256 = 525,336,576
total = 8,030,261,248Every preset on this page reproduces, to the parameter, the total that Hugging Face reports for that checkpoint.
Why do GPT-2 Small’s numbers say 117M, 124M and 137M?
These are three different counts of the same model.
- 117M is the figure OpenAI’s GPT-2 paper printed. It is widely believed to be a miscount; nobody who has loaded the weights gets it.
- 124M (124,439,808) is what
sum(p.numel() for p in model.parameters())returns: the trainable weights. This calculator counts the same thing. - 137M is what the Hugging Face model page shows. The checkpoint also stores a 1,024 × 1,024 causal mask in every layer. That is a buffer, not something the model learns, but the file header counts it: 12 × 1,048,576 = 12,582,912 extra.
nanoGPT reports 123.65M because its get_num_params() subtracts the position table by default. Llama checkpoints carry a similar buffer, RoPE’s inverse frequencies: Llama 2 7B’s file lists 2,048 of them next to its 6,738,415,616 real parameters.
Which parts of a model are embeddings, and why does it matter?
In a small model with a big vocabulary, the embeddings are a large share of the total. In Qwen 2.5 0.5B, the 151,936 × 896 embedding is 136 million of the 494 million parameters, 28% of the model. In Llama 3.1 70B the same kind of matrix is under 2%.
That is why scaling-law papers count non-embedding parameters: the embedding table is a lookup, not compute. It adds almost nothing to the FLOPs per token. Compute and loss track the layers. Memory does not make that distinction: every parameter costs 2 bytes at BF16, embedding or not. That is why the calculator shows the weight footprint from the full total.
How much memory do the parameters take?
Parameters × bytes per parameter: 4 for FP32, 2 for FP16 or BF16, 1 for INT8, about 0.5 for 4-bit. Llama 3.1 8B’s 8.03 billion parameters are 16.06 GB, or 14.96 GiB, at BF16. That is the weights only. Running the model adds a KV cache and framework overhead, and training adds gradients and optimizer state that multiply the figure several times over. Those belong to the VRAM calculators, not this one.
What this does not include
- Mixture-of-experts. An MoE layer holds several FFNs and a router. You can approximate the total by multiplying the FFN size by the expert count, but the calculator does not model routers or shared experts, and the active-parameter count differs from the total.
- Encoder models and encoder–decoder models. BERT adds token-type embeddings, a LayerNorm on the embeddings and a pooler. T5 adds cross-attention in every decoder block. Neither is modelled.
- Latent attention and extra norms. DeepSeek’s multi-head latent attention replaces the K and V projections with low-rank factors, and Gemma 2 adds two more norms per block. Both change the count in ways the fields above cannot express. Pythia’s parallel residual only rearranges the block, so it counts the same as GPT-2.
- Unusual bias layouts. The three bias modes cover GPT-2, Llama/Mistral and Qwen. A model with a bias on the output projection but not on Q, K and V needs a custom sum.
- Buffers. Causal masks and RoPE frequency tables are not parameters, and they are not counted, even though some checkpoints store them.