AI & LLM Calculators
A 7B model needs about 14 GB of VRAM at FP16 and about 4 GB at 4-bit, before the KV cache. These calculators cover the rest of what has to be sized before you build on it: context length and fine-tuning, training compute, tokens per second, vector database storage, chunk sizes and retrieval scores. Every formula is on the page with the numbers that feed it — a model config, a GPU datasheet, a published benchmark — so a result can be checked against the hardware you actually have.
Fine-Tuning VRAM Calculator
Llama 3.1 8B fine-tunes in 7.3 GiB with QLoRA, 19.6 GiB with LoRA and 101 GiB in full. Pick the model, method, LoRA rank, optimizer and batch, and see the four memory terms — plus how many cards of a given size the job needs.
AI & LLM EngineeringKV Cache Calculator
Llama 3.1 8B caches 128 KiB per token — exactly 1 GiB at 8K context, 16 GiB at 128K. Enter the model’s config.json numbers and see the cache per token, per sequence and in total, plus how many sequences fit your memory.
AI & LLM EngineeringLLM VRAM Calculator
Llama 3.1 8B at BF16 needs about 19 GiB with an 8K context — not the 16 GB its download suggests. See weights, KV cache and overhead separately, and which cards fit.
AI & LLM EngineeringTransformer Parameter Calculator
GPT-2 Small has exactly 124,439,808 parameters; Llama 3.1 8B has 8,030,261,248. Enter a model’s config.json numbers to get the exact count, layer by layer, next to the 12 · L · d² shortcut and the memory the weights take.