On this page

To run a large language model, you need enough VRAM for its weights (parameters times bytes per weight), plus the KV cache for your context length, plus 10 to 20% headroom. An 8B model needs about
This guide explains each part of that sum, gives a table for common model sizes and maps them to real GPUs, from a
The formula: weights, cache and headroom
1. Weights: parameters times bytes per weight
Every parameter is stored as a number, and the precision decides its size. NVIDIA's own example: a 7-billion-parameter model loaded in
| Precision | Bytes per weight | Notes |
|---|---|---|
| FP32 | 4 | Training master weights; rarely used for inference |
| FP16 / BF16 | 2 | The usual "full precision" for inference |
| About 1 | ||
| 0.5 in theory, 0.55 to 0.6 in practice | Scales and metadata add a little |
Real files confirm the rule. In llama.cpp's own measurements on
2. KV cache: the memory your context uses
While generating, the model keeps the keys and values of every previous token for every layer. NVIDIA gives the size per token as 2 × layers × (heads × head dimension) × bytes per value. Modern models use
2 × 32 layers × 8 KV heads × 128 × 2 bytes = 131,072 bytes (128 KiB) per token
| Model | KV cache per token (FP16) | 8K context | 32K context | 128K context |
|---|---|---|---|---|
| Needs context extension | ||||
| Needs context extension | ||||
The Qwen3 figures use the published model configurations (14B:
3. Headroom for the runtime
The CUDA context, activations, temporary buffers and memory fragmentation all take space. Plan for 10 to 20% on top of weights plus cache. Some serving engines also reserve most of the free memory for the cache as soon as they start, so the "used" figure in monitoring tools says little about what the model really needs.
VRAM by model size
The table adds weights, an
| Model class (example) | FP16 / BF16 | ||
|---|---|---|---|
| 7B to 8B ( | About | About | About |
| 13B to 14B ( | About | About | About |
| 32B ( | About | About | About |
| 70B ( | About | About | About |
| About | About | About | |
| Not applicable | Not applicable | Fits one |
Mixture-of-experts models break the simple link between size and memory in one direction only.
Which GPUs fit
Memory sizes and bandwidth figures come from NVIDIA's specifications:
| GPU | VRAM | Memory bandwidth | Comfortable fits (8K context) | OffshoreServ |
|---|---|---|---|---|
| 8B at | $61.99 | |||
| 8B at FP16, 14B at | $84.99 | |||
| 32B at | $134.99 | |||
| 32B at | $306.99 | |||
| L40S | 32B at | $418.99 | ||
| 70B at | $355.99 | |||
| Same models as the A100, faster | $581.99 | |||
| 70B at | $558.99 | |||
| 70B at | $1,096.99 | |||
| 70B and 123B at FP16 | $2,348.99 |
Two borderline cases are worth knowing. A
Speed: memory bandwidth sets the pace
Generating a token is, in NVIDIA's words, a
maximum tokens per second ≈ memory bandwidth ÷ size of the weights
| Model and precision | Weights | GPU | Theoretical ceiling |
|---|---|---|---|
| 8B at | About | ||
| 32B at | About | ||
| 70B at | About | ||
| 70B at | H100 SXM5 | About |
Real
Three effects change the picture. Mixture-of-experts models read only their active experts per token, so
Splitting a model across GPUs
When a model does not fit one card, inference engines can split it: by layers across GPUs, or by slicing each layer across GPUs (tensor parallelism). Two cards are not one card with double the memory. Each GPU needs its own runtime overhead, and the cards exchange data at every step.
The link between cards therefore matters. The
Fine-tuning needs far more memory
Everything above is about inference. Full
How to check a specific model in five minutes
- Look at the weight files. The download size of the exact quantized version you plan to run is the best estimate of the weights in memory.
- Open the model's
. Noteconfig.json ,num_hidden_layers andnum_key_value_heads (or hidden size divided by attention heads).head_dim - Compute the cache per token with the formula above, then multiply by your longest expected context and by the number of users you will serve at once.
- Add 10 to 20% and compare with the card's VRAM. If the result is within a few percent of the limit, pick the next size up or a smaller quantization.
- Test with your real prompts. Run the longest document you expect to process and watch memory while it generates.
Quick picks
- Chat assistant or coding helper on an 8B model:
RTX 4090 at full precision, or anRTX A4000 at8-bit . - 14B model with long documents:
RTX 5090 at8-bit . - 32B model:
RTX 5090 at4-bit , or a48 GB card at8-bit for better quality. - 70B model: A100 or H100 at
4-bit ; two H100s at8-bit . 120B-class : one80 GB GPU forgpt-oss-120b ; two to four H100s for dense models such as Mistral Large 2.- Many concurrent users: size the KV cache first, then choose the GPU with the most bandwidth you can afford.
Running your model on OffshoreServ
Our GPU servers are dedicated NVIDIA GPUs, from the
Frequently asked questions
Can a 70B model run on a 24 GB GPU?
Not at a useful speed. At
How much VRAM does an 8B model need?
About
How do I run a model once I have the GPU?
Ollama is the quickest start and vLLM serves an OpenAI-compatible API under load. Our guide on
Should I rent a GPU by the month or by the hour?
By the month if the GPU is busy most days, by the hour for short experiments. Divide the monthly price by the hourly rate you would pay elsewhere to find your
Offshore VPS, dedicated, RDP and GPU servers in seven jurisdictions.


