Free tool
LLM VRAM Calculator
Estimate how much GPU memory a large language model needs for inference. Enter the total parameter count and weight precision, then add your own overhead and KV-cache assumptions.
Raw weights
14.90 GiB
Planning total
17.88 GiB
Min. GPUs (memory only)
1
Planning total = weights 14.90 + overhead 2.98 + KV cache 0.00 GiB. The GPU count is total ÷ usable memory per GPU, rounded up — not a guarantee the model fits or runs.
Weights by precision
| Precision | Bytes / parameter | Raw weights (GiB) |
|---|---|---|
| 32-bit | 4 | 29.80 |
| 16-bit | 2 | 14.90 |
| 8-bit | 1 | 7.45 |
| 4-bit | 0.5 | 3.73 |
The formula
weights GiB = parameters × 10⁹ × bits ÷ 8 ÷ 2³⁰
Worked example: 8B parameters at 16-bit
- 8 × 10⁹ × 16 ÷ 8 = 16,000,000,000 bytes
- 16,000,000,000 ÷ 1,073,741,824 ≈ 14.90 GiB raw weights
- + 20% overhead assumption ≈ 2.98 GiB → ≈ 17.88 GiB with 0 GiB KV cache
- ÷ 24 GiB usable per GPU → memory-only minimum of 1 GPU
Background reading from Hugging Face: KV cache strategies and LLM inference optimization. Compare GPU memory sizes in our GPU database.
FAQ
- How much VRAM do the weights of an 8B model need at 16-bit?
- 8 billion parameters × 16 bits ÷ 8 = 16,000,000,000 bytes, which is about 14.90 GiB for the weights alone. Runtime overhead and KV cache come on top.
- Is 20% overhead a rule?
- No. 20% is only this page's default planning assumption. Real overhead depends on the inference framework, kernels, quantization format and settings, so replace it with a figure you have measured or that your runtime documents.
- Why is the KV cache 0 by default?
- KV-cache memory grows with context length, batch size and model architecture, so it cannot be inferred from parameter count alone. It is excluded unless you enter a value.
- Does the minimum GPU count mean the model will run?
- No. It divides total memory by usable memory per GPU. VRAM is not automatically pooled across GPUs: splitting a model needs tensor or pipeline parallelism support in your framework, which adds its own memory and communication cost.
- What about mixture-of-experts (MoE) models?
- Use the total parameter count, not the active parameters per token. All expert weights normally need to be resident in memory even if only some are used for each token.
- Does this cover training or fine-tuning?
- No. This is an inference estimate. Training also needs memory for gradients, optimizer states and activations, which is typically several times more than the weights.
Estimate only. This is a planning estimate for inference, not a compatibility or performance claim. The runtime and framework, quantization metadata, activations, context length, batch size and parallelism can all require more memory. VRAM is not automatically pooled across GPUs.
Cite this page
“LLM VRAM Calculator.” GPU Data Hub. https://gpudatahub.com/tools/llm-vram-calculator
You are welcome to reference and link to this page. Please link to the URL above.