Launch pricing: every plan costs 30% less than the cheapest offshore competitor we track. See the benchmarkEvery plan 30% under the cheapest offshore host

AI & GPUs

How much VRAM do you need to run an LLM? (2026 guide)

A practical way to size GPU memory for large language models: weights by precision, KV cache for context, headroom, and which GPUs fit models from 8B to 120B parameters.

9 min readBy the OffshoreServ team

Key takeaways

  • Weights: parameters times bytes per weight. 2 bytes at FP16 or BF16, about 1 byte at 8-bit, about 0.5 to 0.6 byte at 4-bit.
  • KV cache: grows with context length and with each concurrent user. About 1 GB per 8K tokens for Llama 3.1 8B, about 2.7 GB for a 70B Llama model.
  • Headroom: keep 10 to 20% free for the runtime, activations and buffers.
  • Speed: generation is limited by memory bandwidth, so a faster-memory GPU produces tokens faster even at equal capacity.
  • Rules of thumb: a 24 GB RTX 4090 runs an 8B model at full precision; a 4-bit 70B model needs about 50 GB, which means one 80 GB GPU or two 32 GB cards.
On this page
  1. The formula: weights, cache and headroom
  2. VRAM by model size
  3. Which GPUs fit
  4. Speed: memory bandwidth sets the pace
  5. Splitting a model across GPUs
  6. Fine-tuning needs far more memory
  7. How to check a specific model in five minutes
  8. Quick picks
  9. Running your model on OffshoreServ
  10. Frequently asked questions

To run a large language model, you need enough VRAM for its weights (parameters times bytes per weight), plus the KV cache for your context length, plus 10 to 20% headroom. An 8B model needs about 16 GB for its weights at FP16, about 8.5 GB at 8-bit and about 5 GB at 4-bit. A 70B model needs about 140 GB, 75 GB and 43 GB respectively, before the cache.

This guide explains each part of that sum, gives a table for common model sizes and maps them to real GPUs, from a 16 GB RTX A4000 to a four-way H100 server.

The formula: weights, cache and headroom

1. Weights: parameters times bytes per weight

Every parameter is stored as a number, and the precision decides its size. NVIDIA's own example: a 7-billion-parameter model loaded in 16-bit precision takes roughly 14 GB. Quantization shrinks that further:

PrecisionBytes per weightNotes
FP324Training master weights; rarely used for inference
FP16 / BF162The usual "full precision" for inference
8-bit (INT8, FP8, Q8)About 1Near-lossless for most models
4-bit (INT4, Q4, MXFP4)0.5 in theory, 0.55 to 0.6 in practiceScales and metadata add a little

Real files confirm the rule. In llama.cpp's own measurements on Llama 3.1 8B, the F16 file is 14.96 GiB, Q8_0 is 7.95 GiB and Q4_K_M is 4.58 GiB, which works out to about 4.9 bits per weight for the popular 4-bit format.

2. KV cache: the memory your context uses

While generating, the model keeps the keys and values of every previous token for every layer. NVIDIA gives the size per token as 2 × layers × (heads × head dimension) × bytes per value. Modern models use grouped-query attention, so the count that matters is the number of key/value heads, which is much smaller than the number of attention heads. Llama 3 8B, for example, has 32 layers and 8 key/value heads of dimension 128:

2 × 32 layers × 8 KV heads × 128 × 2 bytes = 131,072 bytes (128 KiB) per token
ModelKV cache per token (FP16)8K context32K context128K context
Llama 3.1 8B128 KiB1.1 GB4.3 GB17.2 GB
Qwen3-14B160 KiB1.3 GB5.4 GBNeeds context extension
Qwen3-32B256 KiB2.1 GB8.6 GBNeeds context extension
Llama 3.3 70B320 KiB2.7 GB10.7 GB42.9 GB

The Qwen3 figures use the published model configurations (14B: 40 layers; 32B: 64 layers; both with 8 KV heads; their default configuration allows 40,960 positions, so 128K needs context extension). Two consequences matter in practice. The cache is per sequence, so ten concurrent users with 8K contexts need ten times the figures above. And long contexts can outgrow the weights: a 4-bit 8B model at its full 128K context needs more memory for the cache than for itself. Several inference engines can store the cache in 8-bit, which halves it at a small quality cost.

3. Headroom for the runtime

The CUDA context, activations, temporary buffers and memory fragmentation all take space. Plan for 10 to 20% on top of weights plus cache. Some serving engines also reserve most of the free memory for the cache as soon as they start, so the "used" figure in monitoring tools says little about what the model really needs.

VRAM by model size

The table adds weights, an 8K-token context for one user and 10% headroom. Treat the numbers as planning figures: quantization formats differ by a few percent, and longer contexts or more users need more.

Model class (example)FP16 / BF168-bit4-bit
7B to 8B (Llama 3.1 8B)About 19 GBAbout 11 GBAbout 7 GB
13B to 14B (Qwen3-14B)About 32 GBAbout 18 GBAbout 11 GB
32B (Qwen3-32B)About 73 GBAbout 40 GBAbout 24 GB
70B (Llama 3.3 70B)About 157 GBAbout 85 GBAbout 50 GB
120B-class dense (Mistral Large 2, 123B)About 274 GBAbout 147 GBAbout 86 GB
120B-class mixture of experts (gpt-oss-120b)Not applicableNot applicableFits one 80 GB GPU (native MXFP4)

Mixture-of-experts models break the simple link between size and memory in one direction only. gpt-oss-120b has 117 billion parameters but activates only 5.1 billion per token; its expert weights ship in 4-bit MXFP4, which is why OpenAI states it runs on a single 80 GB GPU. All experts still have to sit in memory, so capacity is set by the total size, while speed follows the active parameters.

Which GPUs fit

Memory sizes and bandwidth figures come from NVIDIA's specifications: RTX A4000, RTX 4090 and RTX 5090, RTX 6000 Ada, L40S, A100 and H100. Prices are our monthly prices for a dedicated GPU server.

GPUVRAMMemory bandwidthComfortable fits (8K context)OffshoreServ per month
RTX A400016 GB GDDR6448 GB/s8B at 8-bit, 14B at 4-bit$61.99
RTX 409024 GB GDDR6X1,008 GB/s8B at FP16, 14B at 8-bit$84.99
RTX 509032 GB GDDR71,792 GB/s32B at 4-bit, 14B at 8-bit with long context$134.99
RTX 6000 Ada48 GB GDDR6960 GB/s32B at 8-bit, 14B at FP16$306.99
L40S48 GB GDDR6864 GB/s32B at 8-bit, 14B at FP16$418.99
A100 80 GB80 GB HBM2e1,935 to 2,039 GB/s (PCIe or SXM)70B at 4-bit, gpt-oss-120b, 32B at FP16 (tight)$355.99
H100 80 GB SXM580 GB HBM33.35 TB/sSame models as the A100, faster$581.99
2 × RTX 509064 GB in total1,792 GB/s per card70B at 4-bit, split across both cards$558.99
2 × H100160 GB in total3.35 TB/s per card70B at 8-bit, 123B at 8-bit (tight)$1,096.99
4 × H100320 GB in total3.35 TB/s per card70B and 123B at FP16$2,348.99

Two borderline cases are worth knowing. A 4-bit 32B model needs about 24 GB, which does not leave room on a 24 GB RTX 4090 once the context grows, so the RTX 5090 is the safer single consumer card for that size. A 4-bit 70B model needs about 50 GB, just over a 48 GB card, so plan for 80 GB or two GPUs.

Speed: memory bandwidth sets the pace

Generating a token is, in NVIDIA's words, a memory-bound operation: the time to move weights and cache from memory dominates, not the arithmetic. For a dense model serving one user, every new token reads all the weights once, which gives a hard ceiling:

maximum tokens per second ≈ memory bandwidth ÷ size of the weights
Model and precisionWeightsGPUTheoretical ceiling
8B at 4-bit4.9 GBRTX 4090About 205 tokens/s
32B at 4-bit19.6 GBRTX 5090About 90 tokens/s
70B at 4-bit42.9 GBA100 80 GB (SXM)About 48 tokens/s
70B at 4-bit42.9 GBH100 SXM5About 78 tokens/s

Real single-user numbers land well below these ceilings, because kernels are never perfectly efficient and the cache must be read too. The ranking holds, though, and it explains why the RTX 6000 Ada, with 48 GB but 960 GB/s, generally generates more slowly than an RTX 5090 on a model that fits both.

Three effects change the picture. Mixture-of-experts models read only their active experts per token, so gpt-oss-120b generates far faster than a dense 120B model would. Processing a long prompt (the "prefill") is limited by compute rather than bandwidth, which favors GPUs with more Tensor Core throughput. And batching many users shares each read of the weights across requests, so total throughput rises with concurrency while each user's cache adds to the memory bill.

Splitting a model across GPUs

When a model does not fit one card, inference engines can split it: by layers across GPUs, or by slicing each layer across GPUs (tensor parallelism). Two cards are not one card with double the memory. Each GPU needs its own runtime overhead, and the cards exchange data at every step.

The link between cards therefore matters. The RTX 4090, RTX 5090, RTX 6000 Ada and L40S have no NVLink, according to NVIDIA's specifications, so they communicate over PCIe. The H100 SXM supports NVLink at up to 900 GB/s. For a split 4-bit 70B model on two RTX 5090s, PCIe is workable; for large tensor-parallel deployments, check the interconnect of the exact server before you plan around it.

Fine-tuning needs far more memory

Everything above is about inference. Full fine-tuning is a different scale: with mixed-precision Adam, weights, gradients and optimizer states take 16 bytes per parameter before activations, so an 8B model needs about 128 GB. Parameter-efficient methods change that: the QLoRA authors fine-tuned a 65B model on a single 48 GB GPU by training small adapters on top of a frozen 4-bit model.

How to check a specific model in five minutes

  1. Look at the weight files. The download size of the exact quantized version you plan to run is the best estimate of the weights in memory.
  2. Open the model's config.json. Note num_hidden_layers, num_key_value_heads and head_dim (or hidden size divided by attention heads).
  3. Compute the cache per token with the formula above, then multiply by your longest expected context and by the number of users you will serve at once.
  4. Add 10 to 20% and compare with the card's VRAM. If the result is within a few percent of the limit, pick the next size up or a smaller quantization.
  5. Test with your real prompts. Run the longest document you expect to process and watch memory while it generates.

Quick picks

  • Chat assistant or coding helper on an 8B model: RTX 4090 at full precision, or an RTX A4000 at 8-bit.
  • 14B model with long documents: RTX 5090 at 8-bit.
  • 32B model: RTX 5090 at 4-bit, or a 48 GB card at 8-bit for better quality.
  • 70B model: A100 or H100 at 4-bit; two H100s at 8-bit.
  • 120B-class: one 80 GB GPU for gpt-oss-120b; two to four H100s for dense models such as Mistral Large 2.
  • Many concurrent users: size the KV cache first, then choose the GPU with the most bandwidth you can afford.

Running your model on OffshoreServ

Our GPU servers are dedicated NVIDIA GPUs, from the RTX A4000 to four-way H100 nodes. The RTX range covers most 8B to 32B work, the data-center range adds 48 GB and 80 GB cards, and multi-GPU servers handle 70B models at higher precision and 120B-class dense models. CUDA drivers come pre-installed on our Linux images, described in the GPU images documentation. Servers are ready in 1 to 24 hours in Moldova, the Netherlands, Iceland or Romania, depending on the model, and are paid in crypto with no identity checks.

Frequently asked questions

Can a 70B model run on a 24 GB GPU?

Not at a useful speed. At 4-bit, a 70B model's weights alone take about 40 GB, so on a 24 GB card most layers would sit in system RAM and generation would slow to a crawl. Use two RTX 5090s, a 48 GB card, or an 80 GB A100 or H100.

How much VRAM does an 8B model need?

About 16 GB at 16-bit, 9 GB at 8-bit and 5 GB at 4-bit for the weights, plus the KV cache for your context. A 24 GB card runs an 8B model at full precision with room for long contexts.

How do I run a model once I have the GPU?

Ollama is the quickest start and vLLM serves an OpenAI-compatible API under load. Our guide on self-hosting an LLM on a GPU server walks through both, including how to keep the endpoint private.

Should I rent a GPU by the month or by the hour?

By the month if the GPU is busy most days, by the hour for short experiments. Divide the monthly price by the hourly rate you would pay elsewhere to find your break-even: we do the math in dedicated GPU vs cloud GPU.

Host it where the law is on your side.

Offshore VPS, dedicated, RDP and GPU servers in seven jurisdictions. No KYC, paid in crypto.

Welcome back

Sign in to manage your servers and your balance.

No KYCHuman check by Cloudflare TurnstileNo tracking