AI & Machine Learning
AI Image Generation VRAM Calculator
Estimate how much GPU memory an image generation workflow needs, and what it takes to run it on a smaller card.
Free to useNo sign-up requiredNo watermarkRuns in your browser
Last reviewed: 30 September 2026 by Vishal Senthilkumar
Image generation memory depends on more than the model. Resolution and batch size drive the working memory, the VAE decode at the end can be the biggest spike, and every LoRA and ControlNet you add sits in VRAM alongside the base model.
Pick a model class, set your resolution and add-ons, and this calculator estimates the memory in use, a minimum with common optimisations, and a recommended size with headroom. Usage varies a great deal between ComfyUI, Automatic1111/Forge and diffusers, and with attention and offloading settings, so treat the result as a planning estimate.
How this tool works
Choose a model class
SD 1.5-class, SDXL-class or a large Flux-class transformer. The preset fills in typical file sizes, which you can edit.
Set resolution, batch and precision
Width, height, images per batch, and FP16, FP8 or 4-bit weights.
Add LoRAs and ControlNets
How many, and how large.
Choose VAE and text encoder handling
Tiled VAE decoding, and whether the text encoder stays in VRAM.
Read the estimate and the table
Estimated, minimum and recommended VRAM, and a verdict for common card sizes.
How it works
Weights are the diffusion model and VAE, the text encoder (if kept in VRAM), any LoRAs and any ControlNets, scaled by precision: FP16 as entered, FP32 twice that, FP8 about half, and 4-bit formats about 0.28×. Class sizes come from published parameter counts - roughly 0.9B for an SD 1.5-class UNet, 2.6B for SDXL-class, and about 12B for a Flux-class transformer, whose T5-XXL text encoder is itself close to 5B parameters.
Denoising activations grow with the image area and the batch. The calculator uses an approximate figure per 1024 × 1024 of image area for each architecture, assuming memory-efficient attention, and adds a quarter for each ControlNet running alongside.
The VAE turns the finished latent into pixels, and decoding a large image in one go can need several gigabytes. Because it runs after denoising, the peak is whichever of the two is larger. Tiled decoding caps it at roughly the cost of a few small tiles.
The estimate is 1 GB of runtime overhead plus weights plus that peak. The minimum assumes the text encoder is offloaded after use, tiled VAE decoding and attention slicing. The recommendation adds 20% headroom for previews, upscalers and other applications using the GPU.
Common use cases
- Checking whether a GPU you own can run SDXL or a Flux-class model.
- Choosing between FP16, FP8 and 4-bit versions of a large model.
- Planning how many LoRAs and ControlNets a workflow can stack.
- Finding the largest batch or resolution before out-of-memory errors.
- Sizing a cloud GPU for an image generation service.
Why the same workflow uses different VRAM in different tools
ComfyUI, Automatic1111/Forge and diffusers manage memory differently. Some load and unload models between steps automatically, some keep everything resident for speed, and most switch strategies when they detect a smaller card. Attention implementations (PyTorch SDPA, xformers) and options such as attention slicing, model CPU offload and sequential offload can change peak memory by several gigabytes.
That is why this tool gives a range - a minimum with optimisations, an estimate for a typical setup and a recommendation with headroom - rather than one number. If a workflow runs out of memory, try tiled VAE decoding, offloading the text encoder or a lower-precision model before assuming you need a larger card.
Formula
Weights in VRAM
(model + text encoder + ControlNets) × precision factor + LoRAs
Denoising activations
GB per 1024² × (width × height ÷ 1024²) × batch × (1 + 0.25 × ControlNets)
Estimate
1 GB + weights + max(activations, VAE decode)
Recommended
estimate × 1.2
Worked examples
SDXL-class at 1024 × 1024
A 5.3 GB model plus 1.6 GB of text encoders at FP16 is 6.9 GB of weights. Denoising needs about 2 GB, but an untiled VAE decode at this size needs about 3 GB, so the estimate is 1 + 6.9 + 3 = 10.9 GB, with 13.08 GB recommended. With the text encoder offloaded and tiled decoding, about 7.3 GB - which is why SDXL is possible, slowly, on 8 GB cards.
SD 1.5-class batch of four with LoRAs and a ControlNet
At 512 × 512, batch 4, two 144 MB LoRAs and one 0.72 GB ControlNet, the weights come to 3.16 GB and activations to 3.75 GB, for an estimate of 7.91 GB and 9.49 GB recommended. With optimisations it drops to about 5.78 GB.
Frequently asked questions
How much VRAM do I need for SDXL?
For a single 1024 × 1024 image at FP16, this estimate is about 11 GB in use and 13 GB recommended, so a 16 GB card is comfortable and a 12 GB card workable. With the text encoder offloaded and tiled VAE decoding, SDXL-class models can run on 8 GB cards, more slowly.
Can I run a Flux-class model on a 12 GB or 16 GB card?
At 16-bit the transformer alone is around 24 GB, so not without help. With offloading of the large T5 text encoder and tiled decoding, an FP8 version comes within reach of a 16 GB card and a 4-bit version of a 12 GB card, at some cost to speed and possibly quality.
Do LoRAs use much VRAM?
Usually little - tens to a few hundred megabytes each - because they are small low-rank updates. Some tools merge them into the model weights, in which case they add almost nothing while generating.
Why is the VAE decode so large?
The VAE upsamples the latent eight times in each direction through convolutional layers, holding large full-resolution feature maps. Tiled decoding processes the image in pieces, which caps the peak.
Is the estimate exact?
No. It is built from approximate model sizes and rules of thumb, and real usage depends on your software, settings and driver. Use it to plan, then check with your tool’s memory readout.
