Developer tools
GPU Memory Calculator
Estimate GPU memory for model weights, gradients, optimizer states, and activations locally in your browser.
Enter a model size in billions of parameters, choose its numeric precision and whether you are training or running inference, and add any activations, KV cache, and runtime memory you expect. Weights, gradients, optimizer states, and FP32 master weights follow the usual byte-per-parameter rules, and 1 GB here means 1024 MB, matching how GPUs are rated. The estimate runs locally in your browser; your values are not uploaded or stored.
Size a model for the GPU before you train or deploy it.
Enter a model size in billions of parameters, its numeric precision, and whether you are training or running inference. Training adds gradients, optimizer states for Adam or AdamW and SGD, and optional FP32 master weights for mixed-precision runs; an optional activations and KV cache entry covers memory the model parameters do not. Enter an optional per-card GPU size to see whether the total fits on one card. Memory uses 1 GB = 1024 MB, matching how GPUs are rated. Everything runs locally in your browser, so your values are not uploaded or stored.
Frequently Asked Questions
Everything you need to know about this tool, how it works, and privacy.
What does this GPU memory calculator estimate?
It estimates the device memory a model needs from the number of parameters in billions, the numeric precision of the weights, and whether the workload is training or inference. The result breaks the total into weights, gradients, optimizer states, optional FP32 master weights, and any activations or other memory you enter yourself.
How is training memory different from inference memory?
Inference mostly needs the model weights at the chosen precision. Training also stores gradients of the same size as the weights, optimizer state: Adam or AdamW keeps two FP32 states per parameter, SGD with momentum keeps one, and plain SGD keeps none. Many mixed-precision runs also keep an FP32 master copy of the weights. Choose the purpose on the tool to include or exclude those pieces.
What should I enter for activations and other memory?
Activations, KV cache, CUDA context, and framework buffers live outside the parameter budget and depend on batch size, sequence length, and your stack. Add your measured or expected figure in GB using 1 GB = 1024 MB, or leave it blank to see only the parameter-driven total.
Do my values leave my browser?
No. The estimate and the copy action run locally in your browser. CodeASystem does not upload or store the model size, precision, or memory figures you enter.