AI & Machine Learning
AI GPU Requirement Calculator (VRAM Estimate)
Estimate how much GPU memory a model needs for inference or fine-tuning, and which class of GPU that points to.
Free to useNo sign-up requiredNo watermarkRuns in your browser
Last reviewed: 30 September 2026 by Vishal Senthilkumar
Whether a model runs on a GPU comes down mostly to memory. The weights have to fit, and so does everything that grows with use: the KV cache for long conversations and many simultaneous users, or the gradients, optimizer state and activations when you fine-tune.
This calculator estimates each part from the model size, precision, context length and batch size, adds runtime overhead and headroom, and compares the result with common GPU memory sizes. It gives a class of GPU to look at, not a single right answer - real usage depends on the software you run.
How this tool works
Choose the workload
Inference, LoRA, QLoRA, full fine-tuning, image generation or another model type.
Enter the model size and precision
Parameters in billions, and whether the weights are FP16, 8-bit or 4-bit.
Enter context and batch
Tokens per sequence and how many sequences run at once; for images, the resolution.
Read the estimate and the table
Estimated and recommended VRAM, the suggested GPU class, and a fit verdict for common sizes.
How it works
Weights take parameters × bytes per parameter: 4 bytes at FP32, 2 at FP16 or BF16, 1 at 8-bit, and about 0.56 at 4-bit once the quantisation scales are included. So an 8-billion-parameter model is about 16 GB at FP16 and about 4.5 GB at 4-bit.
For inference, the KV cache stores attention keys and values for every token in every sequence: 2 × layers × KV heads × head dimension × bytes per value, per token. Modern models with grouped-query attention need roughly 100-350 KB per token at 16-bit; the calculator estimates it from model size, and you can enter the exact figure from the model’s configuration. Multiply by context length and batch size.
For fine-tuning, full training with Adam in mixed precision needs about 16 bytes per parameter - weights, gradients, FP32 master weights and two optimizer moments - before activations. LoRA freezes the base model and trains small adapters (assumed here to be 0.5% of parameters); QLoRA does the same over a 4-bit base. Activations are estimated assuming gradient checkpointing.
A runtime overhead of 10% plus 1 GB per process covers the CUDA context, framework buffers and fragmentation. The recommendation adds 20% headroom on top, so longer inputs or spikes do not cause out-of-memory errors.
Common use cases
- Checking whether an open-weight model will run on a GPU you already have.
- Choosing between 16-bit and 4-bit weights for local inference.
- Seeing how context length and concurrent users grow the KV cache.
- Deciding between LoRA, QLoRA and full fine-tuning for a given budget.
- Sizing cloud GPU instances before you rent them.
Ways to fit a model into less memory
- Quantise the weights to 8-bit or 4-bit - the biggest single reduction, with some loss of quality at 4-bit.
- Shorten the context or reduce the batch size to shrink the KV cache.
- Quantise the KV cache to 8-bit, where your inference server supports it.
- For training, use LoRA or QLoRA, gradient checkpointing, 8-bit optimizers or optimizer sharding across GPUs.
- Offload part of the model to CPU memory - it runs, but much more slowly.
Formula
Weights
parameters × bytes per parameter
KV cache
2 × layers × KV heads × head dim × bytes × context × batch
Full fine-tune state
≈ 16 bytes × parameters (+ activations)
Adam optimizer with mixed precision and no sharding.
Estimate
(weights + cache/activations + training state) × 1.1 + 1 GB, per job
Recommended
estimate × 1.2
Worked examples
Serving an 8B model at FP16
Weights are 8 × 2 = 16 GB. With a KV cache of 128 KB per token and an 8,192-token context, the cache is about 1.05 GB. Adding 10% and 1 GB gives an estimate of 19.75 GB and a recommendation of 23.70 GB - the 24 GB class, with 16 GB cards not fitting.
Full fine-tuning a 7B model
At 16 bytes per parameter, weights plus training state are 14 + 98 = 112 GB before activations. With overhead that is about 124 GB, and about 149 GB recommended - more than one 80 GB card, so two with the model sharded, or a switch to LoRA or QLoRA.
Frequently asked questions
How accurate is this estimate?
It is a planning estimate. Real memory use depends on the architecture, the inference or training framework, attention kernels and settings such as how much KV cache is pre-allocated. Some servers deliberately reserve most of the GPU up front. Test on your own setup before committing to hardware.
Why does 4-bit use more than half a byte per parameter?
Quantised formats store scaling factors alongside the 4-bit values, and some layers are often kept at higher precision. Around 4.5 bits per parameter is a reasonable average.
Can I split a model across two GPUs?
Yes - most inference and training frameworks can split a model across GPUs, and the total memory is roughly shared between them, with some duplication and communication overhead. The calculator shows how many 80 GB cards the recommendation would need if split.
Which GPU should I buy?
This tool suggests a memory class rather than a product, because memory is only part of the decision: speed, memory bandwidth, software support, power and price matter too. Use the class to narrow the options.
Why do my measured figures differ from the estimate?
Common reasons are an older model without grouped-query attention (a much larger KV cache), a framework that pre-allocates memory, or extra processes on the same GPU. Enter the exact KV cache per token from the model’s configuration for a closer inference figure.
