How much VRAM does your LLM need?
Pick a model, set the context and the number of concurrent users, and get the GPU count — with the real KV cache of today’s architectures: sliding windows, hybrid linear attention, MLA, and compressed attention. Verified presets, or paste any Hugging Face link.
Frequently asked questions
How much VRAM do I need to run a 70B model?
Take Llama 3.3 70B. Its weights need 131 GiB in BF16 (2 × H100 80 GB), 65.7 GiB in FP8 and about 39.8 GiB in 4-bit (1 × H100, or 2 × RTX 4090). Then add the KV cache of each request: 2.5 GiB at 8K tokens and 10.0 GiB at 32K.
Can I run an LLM on a 24 GB GPU like the RTX 4090?
Yes, models up to about 35B parameters fit in 4-bit with an 8K-token context: Llama 3.1 8B, Mistral Small 3.2 24B, Qwen3 32B, Qwen3.5 4B, Qwen3.5 9B, Qwen3.8 27B, Qwen3.6 35B-A3B (MoE), Gemma 4 12B, Gemma 4 31B, Gemma 4 26B-A4B (MoE), gpt-oss 20B (MoE), Nemotron 3 Nano 30B-A3B, Qwen3-Coder 30B-A3B (MoE). Larger models need several GPUs or a card with more memory.
How many GPUs do I need to run an LLM?
Each GPU must hold its share of the weights plus the KV cache of every concurrent request, under the memory the serving engine allows itself (90% by default in vLLM). The calculator tries 1, 2, 4, 8 GPUs and more, splitting the model with tensor parallelism inside a server and pipeline parallelism across servers, and returns the fewest that fit.
Why is VRAM higher than the model size on disk?
The file holds only the weights. At serving time you also pay for the KV cache, which grows with context × concurrent sequences, runtime buffers such as CUDA graphs and activation workspace, and the share of memory the engine leaves free on purpose.
Why does the KV cache differ so much between models of the same size?
2026 architectures cache very different amounts per token. Qwen 3.5+ keeps a KV cache on only 1 layer in 4, Gemma 4 and gpt-oss limit most layers to a short sliding window, MLA models store one compressed latent instead of keys and values, and DeepSeek V4 compresses the sequence itself. At 128K context, Llama 3.3 70B needs 40 GiB of KV per sequence; Qwen3.8 27B needs 8 GiB.
Can quantization make a model fit on fewer GPUs?
Yes. Weights take parameters × bits per weight, so going from BF16 to 8 bits halves them and 4-bit formats cut them by about 3.5×. As a rule of thumb, 8-bit is near lossless, 5 to 6-bit loses very little and 4-bit loses a little more, depending on the model. Set a fixed number of GPUs and the calculator shows the most accurate format that fits.
Does quantizing the weights shrink the KV cache?
No. The weight format (FP8, NVFP4, GGUF Q4_K_M…) and the KV cache precision are set independently. An FP8 KV cache halves the cache whatever the weights are stored in, at a small accuracy cost.
How much memory does fine-tuning take?
Full fine-tuning with AdamW in mixed precision costs about 16 bytes per parameter before activations: BF16 weights and gradients, an FP32 master copy and two FP32 optimizer moments. LoRA keeps the frozen model at 2 bytes per parameter, QLoRA at about 0.52. Activations then scale with sequence length × batch, not with parameter count.