Fine-tuning Llama 3.1 8B takes about 7.3 GiB of VRAM with QLoRA, 19.6 GiB with LoRA and 101 GiB for full BF16 fine-tuning — at 2,048 tokens, batch 1, gradient checkpointing on and 32-bit AdamW. The full figure is 12 bytes per parameter before activations: 2 for the weight, 2 for its gradient, 8 for the optimizer’s two moments. LoRA and QLoRA pay those 10 bytes only on the adapter, which is under 1% of the model.
The four things that take memory
Training memory is four tensors of known shape plus a margin. With P the model’s parameters, T the trainable ones, L layers, h the hidden size, V the vocabulary, s the sequence length and b the per-GPU micro-batch:
weights = P × bytes per weight 4 FP32 · 2 BF16 · 0.5 NF4 (QLoRA base)
gradients = T × bytes per gradient 2 BF16 · 4 FP32
optimizer = T × bytes per state 8 AdamW · 2 AdamW 8-bit · 12 with FP32 master copy
activations = L × s × b × h × k + s × b × V × 8
total = (weights + gradients + optimizer + activations) × 1.1- k is 2 with gradient checkpointing (one BF16 hidden state kept per layer) and 34 without it, the per-layer figure from Korthikanti et al. 2022, Reducing Activation Recomputation in Large Transformer Models, with FlashAttention removing the attention-score term.
- s × b × V × 8 is the logits: upcast to FP32 for the loss, plus a gradient of the same shape.
- 1.1 covers the CUDA context, cuBLAS workspace and allocator fragmentation. It is the one term that is an estimate rather than a tensor; everything else is exact.
For full fine-tuning T = P. For LoRA and QLoRA the base is frozen and T is the adapter, which the tool computes from the model’s own dimensions instead of guessing a percentage. An adapter of rank r on a projection of shape in × out adds r × (in + out) parameters; summed over a Llama-style block with k/v width kv and intermediate size i:
attention only T = L × r × (6h + 2kv)
all linear T = L × r × (9h + 2kv + 3i)That is what PEFT’s print_trainable_parameters() reports, not an approximation: Llama 3 8B at r = 16 on all linear layers gives 41,943,040, the line at the top of every Llama 3 LoRA notebook.
The formula has no zero case to hide — but it has a fit case that the naive version gets wrong. Across several GPUs with FSDP or ZeRO-3, weights, gradients and optimizer states shard evenly while every card keeps the activations of its own micro-batch. The per-card figure is static ÷ n + activations, so a micro-batch whose activations alone exceed one card does not fit on any number of them; the tool says so rather than dividing the total by the card size.
A worked example, step by step
Llama 3.1 8B — 8,030,261,248 parameters, 32 layers, hidden size 4,096, intermediate 14,336, k/v width 1,024, vocabulary 128,256. QLoRA, r = 16 on all linear layers, 32-bit AdamW, 2,048 tokens, batch 1, gradient checkpointing.
weights = 8,030,261,248 × 0.5 = 4,015,130,624 B = 3.74 GiB
T = 32 × 16 × (9×4096 + 2×1024 + 3×14336) = 41,943,040 (0.52%)
gradients = 41,943,040 × 2 = 83,886,080 B = 0.08 GiB
optimizer = 41,943,040 × 8 = 335,544,320 B = 0.31 GiB
activations = 32 × 2048 × 4096 × 2 = 536,870,912 B = 0.50 GiB
+ 2048 × 128,256 × 8 = 2,101,346,304 B = 1.96 GiB
subtotal = 6.59 GiB
total = 6.59 × 1.1 = 7.25 GiB7.25 GiB is 7.78 GB in decimal. It fits a 12 GB card with room for a longer sequence, and a 24 GB card with room for a micro-batch of 4.
The same model, full BF16 fine-tuning: the static term is 8.03 billion × 12 bytes = 96.36 GB = 89.75 GiB, the activations are unchanged at 2.46 GiB, and the total is (89.75 + 2.46) × 1.1 = 101.4 GiB. On 80 GB cards that is 2 with FSDP: 44.9 GiB of shards plus 2.46 GiB of activations each.
LoRA sits between them at 19.6 GiB: the same 41.9M-parameter adapter as QLoRA, but a BF16 base of 14.96 GiB instead of 3.74.
How much VRAM does a 7B, 14B, 32B or 70B model need?
All rows: 2,048 tokens, micro-batch 1, gradient checkpointing, r = 16 on all linear layers, 32-bit AdamW. The GPU counts assume FSDP or ZeRO-3 and 92% of each card usable.
| Model | Full (BF16) | LoRA | QLoRA |
|---|---|---|---|
| Mistral 7B | 90.1 GiB · 2 × 80 GB | 16.4 GiB · 1 × 24 GB | 5.2 GiB · 1 × 8 GB |
| Qwen 2.5 7B | 96.6 GiB · 2 × 80 GB | 19.0 GiB · 1 × 24 GB | 7.3 GiB · 1 × 12 GB |
| Llama 3.1 8B | 101.4 GiB · 2 × 80 GB | 19.6 GiB · 1 × 24 GB | 7.3 GiB · 1 × 12 GB |
| Gemma 2 9B | 118.5 GiB · 2 × 80 GB | 24.4 GiB · 1 × 32 GB | 10.2 GiB · 1 × 12 GB |
| Phi-4 14B | 182.8 GiB · 3 × 80 GB | 33.3 GiB · 1 × 48 GB | 10.7 GiB · 1 × 12 GB |
| Qwen 2.5 32B | 406.7 GiB · 6 × 80 GB | 72.4 GiB · 2 × 48 GB | 22.1 GiB · 1 × 32 GB |
| Llama 3.1 70B | 872.3 GiB · 13 × 80 GB | 151.6 GiB · 4 × 48 GB | 43.2 GiB · 1 × 48 GB |
Two things the table shows that a per-billion rule of thumb hides. Mistral 7B and Llama 3.1 8B are the same architecture at the same width, yet QLoRA on Llama costs 2 GiB more — that is the 128K vocabulary against Mistral’s 32K, paid in the logits. And a 70B QLoRA at 43 GiB is a 48 GB card with almost no headroom. The QLoRA paper’s headline run — 65B on a single 48 GB GPU — used 512-token sequences; the tool puts the same job at 42.3 GiB there and at 44.8 GiB with 2,048 tokens, which no longer fits.
Why is the activation line bigger than the adapter?
Because LoRA does not touch it. The adapter shrinks gradients and optimizer states — the terms that scale with T — but the backward pass still runs through every frozen layer to reach the adapters underneath, so it saves exactly the tensors full fine-tuning would. In the worked example the adapter’s gradients and optimizer states together are 0.39 GiB; the activations are 2.46 GiB.
Of that, 1.96 GiB is the last layer. Llama 3’s 128,256-token vocabulary makes the logits tensor 2,048 × 128,256 elements, held in FP32 for the loss and again for its gradient. On Gemma 2’s 256,000-token vocabulary it is 3.9 GiB at the same sequence length; on Mistral’s 32,000 it is 0.5 GiB. Frameworks that fuse the loss with the LM head (Unsloth’s cut cross-entropy, Liger’s fused linear cross-entropy) never materialise this tensor, and for small-vocabulary models their gain is small; for Llama 3 and Gemma it is the largest single saving available after checkpointing.
The rest scales with s × b, so doubling the sequence length or the micro-batch doubles it, and both are the same thing to the memory: 8,192 tokens at batch 1 and 2,048 at batch 4 are both 15.4 GiB in the worked example. Gradient accumulation, by contrast, costs nothing — it reaches the same effective batch by running the small micro-batch several times before the optimizer steps.
Turning gradient checkpointing off multiplies the per-layer term by 17 — 0.50 GiB becomes 8.50, and the QLoRA run becomes 16.1 GiB. It buys back roughly a third of the training time; on a card where the job already fits, that is the trade to consider.
Is it 12 or 16 bytes per parameter?
Both figures circulate and both are right about different setups. The QLoRA paper’s “more than 780 GB” for full 16-bit fine-tuning of a 65B model is 12 bytes per parameter: a BF16 weight (2), its BF16 gradient (2) and AdamW’s two FP32 moments (8). That is pure-BF16 training: the model loaded with torch_dtype=torch.bfloat16 and trained as it is, which is how every QLoRA and most LoRA recipes run.
The 16 comes from the ZeRO paper’s mixed-precision accounting, where the optimizer also keeps an FP32 master copy of every weight and updates that, casting down to BF16 for the forward pass — 4 more bytes per trainable parameter, 12 in the optimizer line. That is what happens when a model loaded in FP32 is trained with bf16=True (autocast), and what DeepSpeed and FSDP mixed precision do by default; it is the more numerically careful choice for long runs. The tool’s third optimizer option models it: full fine-tuning of the 8B goes from 101 to 134 GiB.
The lever in the other direction is bitsandbytes’ 8-bit AdamW, which stores the two moments block-wise in one byte each. On the adapter of a QLoRA run it saves a quarter of a gigabyte; on full fine-tuning of the 8B it takes the total from 101 GiB to 52 — one 80 GB card instead of two.
What this does not include
- NF4 quantisation constants. QLoRA’s double-quantised base is about 0.52 bytes per parameter, not 0.50; the tool rounds down by ~4% on the weights line of a QLoRA run.
- Embeddings and the LM head as LoRA targets. The adapter formula covers the seven projections of each block, which is the PEFT default. Adding
embed_tokensandlm_headtomodules_to_savetrains them in full — for Llama 3.1 8B that is another 1.05 billion parameters at 10 bytes each. - Mixture-of-experts and multi-head latent attention. A LoRA over Mixtral’s experts or DeepSeek’s compressed k/v does not follow the block formula above, so neither is offered as a preset.
- Sliding-window and hybrid architectures. Gemma 2’s local layers and any Mamba-style block save different activations from a plain transformer layer.
- The 34 × sbh figure without checkpointing is for a standard block with a 4h feed-forward at 16-bit. Llama’s SwiGLU with three matrices at 3.5h is close, not identical; treat that mode as an estimate to within about 15%.
- Paged optimizers, CPU offload, ZeRO-Offload. These move optimizer states out of VRAM rather than shrinking them. The optimizer line is what leaves the card.
- Evaluation and generation during training, which briefly adds a KV cache. For that figure use the KV cache calculator; for whether the finished model will run on a given card, the LLM VRAM calculator.