DeepLearn Tools
Menu
AI & LLM Engineering

Fine-Tuning VRAM Calculator

Llama 3.1 8B fine-tunes in 7.3 GiB with QLoRA, 19.6 GiB with LoRA and 101 GiB in full. Pick the model, method, LoRA rank, optimizer and batch, and see the four memory terms — plus how many cards of a given size the job needs.

By Seifeur Guizeni, Sr. Agentic AI Architect & AI/ML Consultant

num_hidden_layers

hidden_size

intermediate_size

num_key_value_heads × head_dim

vocab_size

Not the effective batch — gradient accumulation costs no memory.

VRAM needed

7.25 GiB

≈ 7.8 GB decimal. Fits on one 24 GB card.

Base weights 4-bit NF4, 8.03B parameters
3.74 GiB
Gradients 41.9M trainable (0.52%)
0.08 GiB
Optimizer states
0.31 GiB
Activations 2,048 tokens per step, incl. FP32 logits
2.46 GiB
CUDA context and buffers estimate, 10%
0.66 GiB

Same job under each method

Full (BF16)
101.42 GiB
LoRA
19.59 GiB
QLoRA
7.25 GiB

Binary units throughout: 1 GiB = 230 bytes, what nvidia-smi reports. Papers quote decimal GB, which reads about 7% larger.

Share this tool

Fine-tuning Llama 3.1 8B takes about 7.3 GiB of VRAM with QLoRA, 19.6 GiB with LoRA and 101 GiB for full BF16 fine-tuning — at 2,048 tokens, batch 1, gradient checkpointing on and 32-bit AdamW. The full figure is 12 bytes per parameter before activations: 2 for the weight, 2 for its gradient, 8 for the optimizer’s two moments. LoRA and QLoRA pay those 10 bytes only on the adapter, which is under 1% of the model.

The four things that take memory

Training memory is four tensors of known shape plus a margin. With P the model’s parameters, T the trainable ones, L layers, h the hidden size, V the vocabulary, s the sequence length and b the per-GPU micro-batch:

weights      = P × bytes per weight       4 FP32 · 2 BF16 · 0.5 NF4 (QLoRA base)
gradients    = T × bytes per gradient     2 BF16 · 4 FP32
optimizer    = T × bytes per state        8 AdamW · 2 AdamW 8-bit · 12 with FP32 master copy
activations  = L × s × b × h × k  +  s × b × V × 8
total        = (weights + gradients + optimizer + activations) × 1.1

For full fine-tuning T = P. For LoRA and QLoRA the base is frozen and T is the adapter, which the tool computes from the model’s own dimensions instead of guessing a percentage. An adapter of rank r on a projection of shape in × out adds r × (in + out) parameters; summed over a Llama-style block with k/v width kv and intermediate size i:

attention only   T = L × r × (6h + 2kv)
all linear       T = L × r × (9h + 2kv + 3i)

That is what PEFT’s print_trainable_parameters() reports, not an approximation: Llama 3 8B at r = 16 on all linear layers gives 41,943,040, the line at the top of every Llama 3 LoRA notebook.

The formula has no zero case to hide — but it has a fit case that the naive version gets wrong. Across several GPUs with FSDP or ZeRO-3, weights, gradients and optimizer states shard evenly while every card keeps the activations of its own micro-batch. The per-card figure is static ÷ n + activations, so a micro-batch whose activations alone exceed one card does not fit on any number of them; the tool says so rather than dividing the total by the card size.

A worked example, step by step

Llama 3.1 8B — 8,030,261,248 parameters, 32 layers, hidden size 4,096, intermediate 14,336, k/v width 1,024, vocabulary 128,256. QLoRA, r = 16 on all linear layers, 32-bit AdamW, 2,048 tokens, batch 1, gradient checkpointing.

weights      = 8,030,261,248 × 0.5                       = 4,015,130,624 B   = 3.74 GiB
T            = 32 × 16 × (9×4096 + 2×1024 + 3×14336)     = 41,943,040   (0.52%)
gradients    = 41,943,040 × 2                            =    83,886,080 B   = 0.08 GiB
optimizer    = 41,943,040 × 8                            =   335,544,320 B   = 0.31 GiB
activations  = 32 × 2048 × 4096 × 2                      =   536,870,912 B   = 0.50 GiB
             + 2048 × 128,256 × 8                        = 2,101,346,304 B   = 1.96 GiB
subtotal                                                                     = 6.59 GiB
total        = 6.59 × 1.1                                                    = 7.25 GiB

7.25 GiB is 7.78 GB in decimal. It fits a 12 GB card with room for a longer sequence, and a 24 GB card with room for a micro-batch of 4.

The same model, full BF16 fine-tuning: the static term is 8.03 billion × 12 bytes = 96.36 GB = 89.75 GiB, the activations are unchanged at 2.46 GiB, and the total is (89.75 + 2.46) × 1.1 = 101.4 GiB. On 80 GB cards that is 2 with FSDP: 44.9 GiB of shards plus 2.46 GiB of activations each.

LoRA sits between them at 19.6 GiB: the same 41.9M-parameter adapter as QLoRA, but a BF16 base of 14.96 GiB instead of 3.74.

How much VRAM does a 7B, 14B, 32B or 70B model need?

All rows: 2,048 tokens, micro-batch 1, gradient checkpointing, r = 16 on all linear layers, 32-bit AdamW. The GPU counts assume FSDP or ZeRO-3 and 92% of each card usable.

Model Full (BF16) LoRA QLoRA
Mistral 7B 90.1 GiB · 2 × 80 GB 16.4 GiB · 1 × 24 GB 5.2 GiB · 1 × 8 GB
Qwen 2.5 7B 96.6 GiB · 2 × 80 GB 19.0 GiB · 1 × 24 GB 7.3 GiB · 1 × 12 GB
Llama 3.1 8B 101.4 GiB · 2 × 80 GB 19.6 GiB · 1 × 24 GB 7.3 GiB · 1 × 12 GB
Gemma 2 9B 118.5 GiB · 2 × 80 GB 24.4 GiB · 1 × 32 GB 10.2 GiB · 1 × 12 GB
Phi-4 14B 182.8 GiB · 3 × 80 GB 33.3 GiB · 1 × 48 GB 10.7 GiB · 1 × 12 GB
Qwen 2.5 32B 406.7 GiB · 6 × 80 GB 72.4 GiB · 2 × 48 GB 22.1 GiB · 1 × 32 GB
Llama 3.1 70B 872.3 GiB · 13 × 80 GB 151.6 GiB · 4 × 48 GB 43.2 GiB · 1 × 48 GB

Two things the table shows that a per-billion rule of thumb hides. Mistral 7B and Llama 3.1 8B are the same architecture at the same width, yet QLoRA on Llama costs 2 GiB more — that is the 128K vocabulary against Mistral’s 32K, paid in the logits. And a 70B QLoRA at 43 GiB is a 48 GB card with almost no headroom. The QLoRA paper’s headline run — 65B on a single 48 GB GPU — used 512-token sequences; the tool puts the same job at 42.3 GiB there and at 44.8 GiB with 2,048 tokens, which no longer fits.

Why is the activation line bigger than the adapter?

Because LoRA does not touch it. The adapter shrinks gradients and optimizer states — the terms that scale with T — but the backward pass still runs through every frozen layer to reach the adapters underneath, so it saves exactly the tensors full fine-tuning would. In the worked example the adapter’s gradients and optimizer states together are 0.39 GiB; the activations are 2.46 GiB.

Of that, 1.96 GiB is the last layer. Llama 3’s 128,256-token vocabulary makes the logits tensor 2,048 × 128,256 elements, held in FP32 for the loss and again for its gradient. On Gemma 2’s 256,000-token vocabulary it is 3.9 GiB at the same sequence length; on Mistral’s 32,000 it is 0.5 GiB. Frameworks that fuse the loss with the LM head (Unsloth’s cut cross-entropy, Liger’s fused linear cross-entropy) never materialise this tensor, and for small-vocabulary models their gain is small; for Llama 3 and Gemma it is the largest single saving available after checkpointing.

The rest scales with s × b, so doubling the sequence length or the micro-batch doubles it, and both are the same thing to the memory: 8,192 tokens at batch 1 and 2,048 at batch 4 are both 15.4 GiB in the worked example. Gradient accumulation, by contrast, costs nothing — it reaches the same effective batch by running the small micro-batch several times before the optimizer steps.

Turning gradient checkpointing off multiplies the per-layer term by 17 — 0.50 GiB becomes 8.50, and the QLoRA run becomes 16.1 GiB. It buys back roughly a third of the training time; on a card where the job already fits, that is the trade to consider.

Is it 12 or 16 bytes per parameter?

Both figures circulate and both are right about different setups. The QLoRA paper’s “more than 780 GB” for full 16-bit fine-tuning of a 65B model is 12 bytes per parameter: a BF16 weight (2), its BF16 gradient (2) and AdamW’s two FP32 moments (8). That is pure-BF16 training: the model loaded with torch_dtype=torch.bfloat16 and trained as it is, which is how every QLoRA and most LoRA recipes run.

The 16 comes from the ZeRO paper’s mixed-precision accounting, where the optimizer also keeps an FP32 master copy of every weight and updates that, casting down to BF16 for the forward pass — 4 more bytes per trainable parameter, 12 in the optimizer line. That is what happens when a model loaded in FP32 is trained with bf16=True (autocast), and what DeepSpeed and FSDP mixed precision do by default; it is the more numerically careful choice for long runs. The tool’s third optimizer option models it: full fine-tuning of the 8B goes from 101 to 134 GiB.

The lever in the other direction is bitsandbytes’ 8-bit AdamW, which stores the two moments block-wise in one byte each. On the adapter of a QLoRA run it saves a quarter of a gigabyte; on full fine-tuning of the 8B it takes the total from 101 GiB to 52 — one 80 GB card instead of two.

What this does not include

Frequently asked questions

How much VRAM do I need to fine-tune a 7B or 8B model?
Llama 3.1 8B at 2,048 tokens with gradient checkpointing: about 7.3 GiB with QLoRA, 19.6 GiB with LoRA and 101 GiB for full BF16 fine-tuning. So a 12 GB card for QLoRA, a 24 GB card for LoRA, and two 80 GB cards for full fine-tuning.
How is fine-tuning VRAM calculated?
Four terms: weights (2 bytes per parameter in BF16, 0.5 for a 4-bit QLoRA base), gradients (2 bytes per trainable parameter), AdamW optimizer states (8 bytes per trainable parameter) and activations, which scale with layers × sequence length × batch × hidden size, plus the FP32 logits. Full fine-tuning trains every parameter, so it costs 12 bytes each before activations.
Does LoRA reduce activation memory?
No. LoRA shrinks the gradient and optimizer terms because only the adapter is trainable, but the backward pass still runs through every frozen layer, so the activations saved are the same as in full fine-tuning. Sequence length, micro-batch and gradient checkpointing are what control that term.
Why does Llama 3 need more VRAM than Mistral 7B for the same QLoRA run?
The vocabulary. The logits tensor is sequence × vocabulary in FP32, twice (loss and gradient). At 2,048 tokens that is 1.96 GiB for Llama 3’s 128,256 tokens against 0.5 GiB for Mistral’s 32,000, and it is the largest single activation in a QLoRA run on Llama 3 or Gemma.
Is it 12 or 16 bytes per parameter for full fine-tuning?
12 for pure BF16 training: 2 weight + 2 gradient + 8 for AdamW’s two FP32 moments — the QLoRA paper’s “780 GB for 65B”. 16 when the optimizer also keeps an FP32 master copy of the weights, as DeepSpeed and FSDP mixed precision do. The optimizer select on the tool switches between them.
How many GPUs do I need for full fine-tuning?
With FSDP or ZeRO-3 the weights, gradients and optimizer states shard evenly across cards, but every card keeps the activations of its own micro-batch. The tool finds the smallest count where static ÷ n + activations fits in 92% of a card — two 80 GB cards for Llama 3.1 8B, thirteen for 70B.