DeepLearn Tools
Menu
AI & LLM Engineering

Transformer Parameter Calculator

GPT-2 Small has exactly 124,439,808 parameters; Llama 3.1 8B has 8,030,261,248. Enter a model’s config.json numbers to get the exact count, layer by layer, next to the 12 · L · d² shortcut and the memory the weights take.

By Seifeur Guizeni, Sr. Agentic AI Architect & AI/ML Consultant

vocab_size

hidden_size / n_embd

num_hidden_layers / n_layer

num_attention_heads

num_key_value_heads — equal to heads for MHA

head_dim, or hidden ÷ heads

intermediate_size — 4 × hidden on GPT-2

Total parameters

8.03B

8,030,261,248 exactly

Token embedding
525,336,576
Per layer attention 41,943,040 · FFN 176,160,768 · norms 8,192
218,112,000
All 32 layers
6,979,584,000
Final norm
4,096
LM head
525,336,576
Non-embedding (Kaplan N)
7.50B

12 · L · d² shortcut: 6.44B

-14.2% against the exact non-embedding count.

Weights alone take

FP32
29.92 GiB
BF16
14.96 GiB
INT8
7.48 GiB
INT4
3.74 GiB

Share this tool

A transformer’s parameter count is the sum of its weight matrices: token embeddings (V × d), each layer’s attention and feed-forward network, the norms, and an output head unless it is tied to the embeddings. GPT-2 Small works out to exactly 124,439,808 parameters; Llama 3.1 8B to 8,030,261,248. The shortcut 12 · L · d² covers the layers alone.

The formula

Take the numbers from the model’s config.json:

One transformer block holds:

attention  = d·(H·h)  +  2·d·(H_kv·h)  +  (H·h)·d        Q, K and V, O
FFN        = 2·d·f   (standard: up, down)
           = 3·d·f   (gated / SwiGLU: gate, up, down)
norms      = 2 × 2d  (LayerNorm: a scale and a shift)
           = 2 × d   (RMSNorm: a scale only)

plus biases where the architecture has them: H·h + 2·H_kv·h on Q, K and V, d on the attention output, and f (or 2f) + d on the FFN.

The whole model is:

total = V·d                   token embedding
      + P·d                   learned position table (0 for RoPE or ALiBi)
      + L × per-block         all layers
      + final norm            d or 2d
      + d·V                   LM head (0 when tied)

The case people get wrong is the tied head. Most small models, GPT-2 and Qwen 2.5 0.5B included, reuse the embedding matrix as the output layer (tie_word_embeddings: true). It is one matrix used twice, so it is counted once. Adding d·V again overstates GPT-2 Small by 38.6 million — about 31%.

Where 12 · L · d² comes from

With plain multi-head attention, Q, K, V and O are each d × d: that is 4d². A standard FFN with f = 4d has two d × 4d matrices: 8d². Together that is 12d² per layer, and Kaplan et al.’s scaling-laws paper (2020) uses N ≈ 12 · L · d² for the non-embedding parameter count.

It is exact for GPT-2-style blocks, apart from biases and norms, and it drifts on modern ones:

For Llama 3.1 8B the shortcut gives 6.44B for the layers; the exact non-embedding count is 7.50B. That is 14% short. For GPT-2 Small it is 0.1% short. The calculator shows both side by side.

A worked example, step by step

GPT-2 Small: V = 50,257, d = 768, L = 12, H = 12, f = 3,072, 1,024 learned positions, LayerNorm, biases on every linear layer, tied head.

token embedding   50,257 × 768                      = 38,597,376
position table     1,024 × 768                      =    786,432

per layer
  attention        4 × 768² + 4 × 768               =  2,362,368
  FFN              2 × 768 × 3,072 + 3,072 + 768    =  4,722,432
  norms            2 × 2 × 768                      =      3,072
                                                      ----------
                                                       7,087,872
all layers        12 × 7,087,872                    = 85,054,464
final norm        2 × 768                           =      1,536
LM head           tied                              =          0
                                                      ----------
total                                               = 124,439,808

That is the “124M” on the model card. Pick the GPT-2 Small preset above and every row matches.

Llama 3.1 8B does the same arithmetic with GQA, a gated FFN, RMSNorm, no biases and an untied head:

token embedding   128,256 × 4,096                        =   525,336,576
per layer
  attention       2 × 4,096² + 2 × 4,096 × 1,024         =    41,943,040
  FFN             3 × 4,096 × 14,336                     =   176,160,768
  norms           2 × 4,096                              =         8,192
all layers        32 × 218,112,000                       = 6,979,584,000
final norm                                               =         4,096
LM head           4,096 × 128,256                        =   525,336,576
total                                                    = 8,030,261,248

Every preset on this page reproduces, to the parameter, the total that Hugging Face reports for that checkpoint.

Why do GPT-2 Small’s numbers say 117M, 124M and 137M?

These are three different counts of the same model.

nanoGPT reports 123.65M because its get_num_params() subtracts the position table by default. Llama checkpoints carry a similar buffer, RoPE’s inverse frequencies: Llama 2 7B’s file lists 2,048 of them next to its 6,738,415,616 real parameters.

Which parts of a model are embeddings, and why does it matter?

In a small model with a big vocabulary, the embeddings are a large share of the total. In Qwen 2.5 0.5B, the 151,936 × 896 embedding is 136 million of the 494 million parameters, 28% of the model. In Llama 3.1 70B the same kind of matrix is under 2%.

That is why scaling-law papers count non-embedding parameters: the embedding table is a lookup, not compute. It adds almost nothing to the FLOPs per token. Compute and loss track the layers. Memory does not make that distinction: every parameter costs 2 bytes at BF16, embedding or not. That is why the calculator shows the weight footprint from the full total.

How much memory do the parameters take?

Parameters × bytes per parameter: 4 for FP32, 2 for FP16 or BF16, 1 for INT8, about 0.5 for 4-bit. Llama 3.1 8B’s 8.03 billion parameters are 16.06 GB, or 14.96 GiB, at BF16. That is the weights only. Running the model adds a KV cache and framework overhead, and training adds gradients and optimizer state that multiply the figure several times over. Those belong to the VRAM calculators, not this one.

What this does not include

Frequently asked questions

How do I calculate the number of parameters in a transformer?
Add the token embedding (vocab × d), L copies of one block, the final norm and the output head. A plain block has 4d² for attention and 8d² for a 4d-wide FFN, which is where 12 · L · d² comes from. Add biases, norms and learned positions for the exact figure, and skip the head if it is tied.
Is 12 · L · d² accurate for modern LLMs?
Not very. It assumes full multi-head attention and a two-matrix FFN that is 4d wide. Grouped-query attention and SwiGLU break both assumptions: for Llama 3.1 8B the shortcut gives 6.44B for the layers against an exact 7.50B, 14% short. For GPT-2 it is within 0.2%.
Why does GPT-2 Small show 117M, 124M or 137M parameters?
124,439,808 is the real trainable count. 117M is the figure printed in the original paper and matches no loaded model. 137M is Hugging Face’s file total, which also counts a 1,024 × 1,024 causal-mask buffer stored in each of the 12 layers.
Do tied embeddings count twice?
No. With tie_word_embeddings the output layer reuses the input embedding matrix, so it is one set of weights counted once. Counting it twice overstates GPT-2 Small by 38.6 million, about 31%. Untied models such as Llama 3 carry two separate vocab × d matrices.
What are non-embedding parameters?
The total minus the token and position embedding tables. Scaling-law papers such as Kaplan et al. use this count because an embedding lookup costs almost no compute. Memory counts every parameter, so weight size is always computed from the full total.
How much memory do the parameters take?
Parameters × bytes each: 2 at BF16, 1 at INT8, about 0.5 at 4-bit. Llama 3.1 8B is 16.06 GB (14.96 GiB) at BF16. That is weights only; inference adds a KV cache, and training adds gradients and optimizer state.