---
title: "How Much VRAM Do You Need to Run an LLM? (2026 Guide)"
description: "How much GPU memory an LLM needs: the parameters × bytes rule, KV cache and headroom, a VRAM table from 8B to 120B models, which GPUs fit."
url: https://offshoreserv.com/blog/how-much-vram-for-llms
lang: en
updated: 2026-09-26
source: HTML page at the url above (canonical); this is its Markdown version
---

[AI & GPUs](https://offshoreserv.com/blog/category/ai-gpus)

# How much VRAM do you need to run an LLM? (2026 guide)

A practical way to size GPU memory for large language models: weights by precision, KV cache for context, headroom, and which GPUs fit models from 8B to 120B parameters.

26 September 2026 9 min read By the OffshoreServ team

Key takeaways

- Weights: parameters times bytes per weight. 2 bytes at FP16 or BF16, about 1 byte at 8-bit, about 0.5 to 0.6 byte at 4-bit.
- KV cache: grows with context length and with each concurrent user. About 1 GB per 8K tokens for Llama 3.1 8B, about 2.7 GB for a 70B Llama model.
- Headroom: keep 10 to 20% free for the runtime, activations and buffers.
- Speed: generation is limited by memory bandwidth, so a faster-memory GPU produces tokens faster even at equal capacity.
- Rules of thumb: a 24 GB RTX 4090 runs an 8B model at full precision; a 4-bit 70B model needs about 50 GB, which means one 80 GB GPU or two 32 GB cards.

To run a large language model, you need enough VRAM for its weights (parameters times bytes per weight), plus the KV cache for your context length, plus 10 to 20% headroom. An 8B model needs about 16 GB for its weights at FP16, about 8.5 GB at 8-bit and about 5 GB at 4-bit. A 70B model needs about 140 GB, 75 GB and 43 GB respectively, before the cache.

This guide explains each part of that sum, gives a table for common model sizes and maps them to real GPUs, from a 16 GB RTX A4000 to a four-way H100 server.

## The formula: weights, cache and headroom

### 1. Weights: parameters times bytes per weight

Every parameter is stored as a number, and the precision decides its size. NVIDIA's own example: a 7-billion-parameter model loaded in 16-bit precision [takes roughly 14 GB](https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/). Quantization shrinks that further:

| Precision | Bytes per weight | Notes |
| --- | --- | --- |
| FP32 | 4 | Training master weights; rarely used for inference |
| FP16 / BF16 | 2 | The usual "full precision" for inference |
| 8-bit (INT8, FP8, Q8) | About 1 | Near-lossless for most models |
| 4-bit (INT4, Q4, MXFP4) | 0.5 in theory, 0.55 to 0.6 in practice | Scales and metadata add a little |

Real files confirm the rule. In llama.cpp's own measurements on Llama 3.1 8B, the [F16 file is 14.96 GiB, Q8_0 is 7.95 GiB and Q4_K_M is 4.58 GiB](https://github.com/ggml-org/llama.cpp/blob/master/tools/quantize/README.md), which works out to about 4.9 bits per weight for the popular 4-bit format.

### 2. KV cache: the memory your context uses

While generating, the model keeps the keys and values of every previous token for every layer. NVIDIA gives the size per token as `2 × layers × (heads × head dimension) × bytes per value`. Modern models use grouped-query attention, so the count that matters is the number of key/value heads, which is much smaller than the number of attention heads. Llama 3 8B, for example, has 32 layers and [8 key/value heads](https://arxiv.org/abs/2407.21783) of dimension 128:

```
2 × 32 layers × 8 KV heads × 128 × 2 bytes = 131,072 bytes (128 KiB) per token
```

| Model | KV cache per token (FP16) | 8K context | 32K context | 128K context |
| --- | --- | --- | --- | --- |
| Llama 3.1 8B | 128 KiB | 1.1 GB | 4.3 GB | 17.2 GB |
| Qwen3-14B | 160 KiB | 1.3 GB | 5.4 GB | Needs context extension |
| Qwen3-32B | 256 KiB | 2.1 GB | 8.6 GB | Needs context extension |
| Llama 3.3 70B | 320 KiB | 2.7 GB | 10.7 GB | 42.9 GB |

The Qwen3 figures use the published model configurations ([14B](https://huggingface.co/Qwen/Qwen3-14B/blob/main/config.json): 40 layers; [32B](https://huggingface.co/Qwen/Qwen3-32B/blob/main/config.json): 64 layers; both with 8 KV heads; their default configuration allows 40,960 positions, so 128K needs context extension). Two consequences matter in practice. The cache is per sequence, so ten concurrent users with 8K contexts need ten times the figures above. And long contexts can outgrow the weights: a 4-bit 8B model at its full 128K context needs more memory for the cache than for itself. Several inference engines can store the cache in 8-bit, which halves it at a small quality cost.

### 3. Headroom for the runtime

The CUDA context, activations, temporary buffers and memory fragmentation all take space. Plan for 10 to 20% on top of weights plus cache. Some serving engines also reserve most of the free memory for the cache as soon as they start, so the "used" figure in monitoring tools says little about what the model really needs.

## VRAM by model size

The table adds weights, an 8K-token context for one user and 10% headroom. Treat the numbers as planning figures: quantization formats differ by a few percent, and longer contexts or more users need more.

| Model class (example) | FP16 / BF16 | 8-bit | 4-bit |
| --- | --- | --- | --- |
| 7B to 8B (Llama 3.1 8B) | About 19 GB | About 11 GB | About 7 GB |
| 13B to 14B (Qwen3-14B) | About 32 GB | About 18 GB | About 11 GB |
| 32B (Qwen3-32B) | About 73 GB | About 40 GB | About 24 GB |
| 70B (Llama 3.3 70B) | About 157 GB | About 85 GB | About 50 GB |
| 120B-class dense ([Mistral Large 2, 123B](https://mistral.ai/news/mistral-large-2407/)) | About 274 GB | About 147 GB | About 86 GB |
| 120B-class mixture of experts ([gpt-oss-120b](https://huggingface.co/openai/gpt-oss-120b)) | Not applicable | Not applicable | Fits one 80 GB GPU (native MXFP4) |

Mixture-of-experts models break the simple link between size and memory in one direction only. gpt-oss-120b has 117 billion parameters but activates only 5.1 billion per token; its expert weights ship in 4-bit MXFP4, which is why OpenAI states it runs on a single 80 GB GPU. All experts still have to sit in memory, so capacity is set by the total size, while speed follows the active parameters.

## Which GPUs fit

Memory sizes and bandwidth figures come from NVIDIA's specifications: [RTX A4000](https://www.nvidia.com/content/dam/en-zz/Solutions/gtcs21/rtx-a4000/nvidia-rtx-a4000-datasheet.pdf), [RTX 4090 and RTX 5090](https://images.nvidia.com/aem-dam/Solutions/geforce/blackwell/nvidia-rtx-blackwell-gpu-architecture.pdf), [RTX 6000 Ada](https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/rtx-6000/proviz-print-rtx6000-datasheet-web-2504660.pdf), [L40S](https://www.nvidia.com/en-us/data-center/l40s/), [A100](https://www.nvidia.com/en-us/data-center/a100/) and [H100](https://www.nvidia.com/en-us/data-center/h100/). Prices are our monthly prices for a dedicated GPU server.

| GPU | VRAM | Memory bandwidth | Comfortable fits (8K context) | OffshoreServ per month |
| --- | --- | --- | --- | --- |
| RTX A4000 | 16 GB GDDR6 | 448 GB/s | 8B at 8-bit, 14B at 4-bit | $61.99 |
| RTX 4090 | 24 GB GDDR6X | 1,008 GB/s | 8B at FP16, 14B at 8-bit | $84.99 |
| RTX 5090 | 32 GB GDDR7 | 1,792 GB/s | 32B at 4-bit, 14B at 8-bit with long context | $134.99 |
| RTX 6000 Ada | 48 GB GDDR6 | 960 GB/s | 32B at 8-bit, 14B at FP16 | $306.99 |
| L40S | 48 GB GDDR6 | 864 GB/s | 32B at 8-bit, 14B at FP16 | $418.99 |
| A100 80 GB | 80 GB HBM2e | 1,935 to 2,039 GB/s (PCIe or SXM) | 70B at 4-bit, gpt-oss-120b, 32B at FP16 (tight) | $355.99 |
| H100 80 GB SXM5 | 80 GB HBM3 | 3.35 TB/s | Same models as the A100, faster | $581.99 |
| 2 × RTX 5090 | 64 GB in total | 1,792 GB/s per card | 70B at 4-bit, split across both cards | $558.99 |
| 2 × H100 | 160 GB in total | 3.35 TB/s per card | 70B at 8-bit, 123B at 8-bit (tight) | $1,096.99 |
| 4 × H100 | 320 GB in total | 3.35 TB/s per card | 70B and 123B at FP16 | $2,348.99 |

Two borderline cases are worth knowing. A 4-bit 32B model needs about 24 GB, which does not leave room on a 24 GB RTX 4090 once the context grows, so the RTX 5090 is the safer single consumer card for that size. A 4-bit 70B model needs about 50 GB, just over a 48 GB card, so plan for 80 GB or two GPUs.

## Speed: memory bandwidth sets the pace

Generating a token is, in NVIDIA's words, [a memory-bound operation](https://developer.nvidia.com/blog/mastering-llm-techniques-inference-optimization/): the time to move weights and cache from memory dominates, not the arithmetic. For a dense model serving one user, every new token reads all the weights once, which gives a hard ceiling:

```
maximum tokens per second ≈ memory bandwidth ÷ size of the weights
```

| Model and precision | Weights | GPU | Theoretical ceiling |
| --- | --- | --- | --- |
| 8B at 4-bit | 4.9 GB | RTX 4090 | About 205 tokens/s |
| 32B at 4-bit | 19.6 GB | RTX 5090 | About 90 tokens/s |
| 70B at 4-bit | 42.9 GB | A100 80 GB (SXM) | About 48 tokens/s |
| 70B at 4-bit | 42.9 GB | H100 SXM5 | About 78 tokens/s |

Real single-user numbers land well below these ceilings, because kernels are never perfectly efficient and the cache must be read too. The ranking holds, though, and it explains why the RTX 6000 Ada, with 48 GB but 960 GB/s, generally generates more slowly than an RTX 5090 on a model that fits both.

Three effects change the picture. Mixture-of-experts models read only their active experts per token, so gpt-oss-120b generates far faster than a dense 120B model would. Processing a long prompt (the "prefill") is limited by compute rather than bandwidth, which favors GPUs with more Tensor Core throughput. And batching many users shares each read of the weights across requests, so total throughput rises with concurrency while each user's cache adds to the memory bill.

## Splitting a model across GPUs

When a model does not fit one card, inference engines can split it: by layers across GPUs, or by slicing each layer across GPUs (tensor parallelism). Two cards are not one card with double the memory. Each GPU needs its own runtime overhead, and the cards exchange data at every step.

The link between cards therefore matters. The RTX 4090, RTX 5090, RTX 6000 Ada and L40S have no NVLink, according to NVIDIA's specifications, so they communicate over PCIe. The H100 SXM supports NVLink at up to 900 GB/s. For a split 4-bit 70B model on two RTX 5090s, PCIe is workable; for large tensor-parallel deployments, check the interconnect of the exact server before you plan around it.

## Fine-tuning needs far more memory

Everything above is about inference. Full fine-tuning is a different scale: with mixed-precision Adam, weights, gradients and optimizer states take [16 bytes per parameter](https://arxiv.org/abs/1910.02054) before activations, so an 8B model needs about 128 GB. Parameter-efficient methods change that: the QLoRA authors [fine-tuned a 65B model on a single 48 GB GPU](https://arxiv.org/abs/2305.14314) by training small adapters on top of a frozen 4-bit model.

## How to check a specific model in five minutes

1. **Look at the weight files.** The download size of the exact quantized version you plan to run is the best estimate of the weights in memory.
2. **Open the model's `config.json`.** Note `num_hidden_layers`, `num_key_value_heads` and `head_dim` (or hidden size divided by attention heads).
3. **Compute the cache per token** with the formula above, then multiply by your longest expected context and by the number of users you will serve at once.
4. **Add 10 to 20%** and compare with the card's VRAM. If the result is within a few percent of the limit, pick the next size up or a smaller quantization.
5. **Test with your real prompts.** Run the longest document you expect to process and watch memory while it generates.

## Quick picks

- **Chat assistant or coding helper on an 8B model:** [RTX 4090](https://offshoreserv.com/offshore-gpu-servers/rtx-4090) at full precision, or an RTX A4000 at 8-bit.
- **14B model with long documents:** [RTX 5090](https://offshoreserv.com/offshore-gpu-servers/rtx-5090) at 8-bit.
- **32B model:** RTX 5090 at 4-bit, or a 48 GB card at 8-bit for better quality.
- **70B model:** [A100](https://offshoreserv.com/offshore-gpu-servers/a100) or [H100](https://offshoreserv.com/offshore-gpu-servers/h100) at 4-bit; two H100s at 8-bit.
- **120B-class:** one 80 GB GPU for gpt-oss-120b; two to four H100s for dense models such as Mistral Large 2.
- **Many concurrent users:** size the KV cache first, then choose the GPU with the most bandwidth you can afford.

## Running your model on OffshoreServ

Our [GPU servers](https://offshoreserv.com/offshore-gpu-servers) are dedicated NVIDIA GPUs, from the RTX A4000 to four-way H100 nodes. The [RTX range](https://offshoreserv.com/offshore-gpu-servers/rtx) covers most 8B to 32B work, the [data-center range](https://offshoreserv.com/offshore-gpu-servers/datacenter) adds 48 GB and 80 GB cards, and [multi-GPU servers](https://offshoreserv.com/offshore-gpu-servers/multi-gpu) handle 70B models at higher precision and 120B-class dense models. CUDA drivers come pre-installed on our Linux images, described in the [GPU images documentation](https://offshoreserv.com/docs/gpu/images). Servers are ready in 1 to 24 hours in Moldova, the Netherlands, Iceland or Romania, depending on the model, and are paid in crypto with no identity checks.

## Frequently asked questions

### Can a 70B model run on a 24 GB GPU?

Not at a useful speed. At 4-bit, a 70B model's weights alone take about 40 GB, so on a 24 GB card most layers would sit in system RAM and generation would slow to a crawl. Use two RTX 5090s, a 48 GB card, or an 80 GB A100 or H100.

### How much VRAM does an 8B model need?

About 16 GB at 16-bit, 9 GB at 8-bit and 5 GB at 4-bit for the weights, plus the KV cache for your context. A 24 GB card runs an 8B model at full precision with room for long contexts.

### How do I run a model once I have the GPU?

Ollama is the quickest start and vLLM serves an OpenAI-compatible API under load. Our guide on [self-hosting an LLM on a GPU server](https://offshoreserv.com/blog/self-host-llm-gpu-server) walks through both, including how to keep the endpoint private.

### Should I rent a GPU by the month or by the hour?

By the month if the GPU is busy most days, by the hour for short experiments. Divide the monthly price by the hourly rate you would pay elsewhere to find your break-even: we do the math in [dedicated GPU vs cloud GPU](https://offshoreserv.com/blog/dedicated-gpu-vs-cloud-gpu).

**Host it where the law is on your side.**

Offshore VPS, dedicated, RDP and GPU servers in seven jurisdictions. No KYC, paid in crypto.

## More from the blog.

- [AI & GPUs Dedicated GPU server vs cloud GPU: cost, privacy and speed Monthly or hourly? How to find the break-even point between a dedicated GPU server and a cloud GPU, and what changes for availability, speed and privacy.26 September 2026 8 min read](https://offshoreserv.com/blog/dedicated-gpu-vs-cloud-gpu)
- [AI & GPUs How to self-host an LLM on a GPU server (Ollama, vLLM) From an empty GPU server to a private, OpenAI-compatible endpoint: which card to rent, Ollama and vLLM setup, TLS and keys, benchmarks, privacy and cost.26 September 2026 10 min read](https://offshoreserv.com/blog/self-host-llm-gpu-server)
- [Law & jurisdictions 5, 9 and 14 Eyes countries: what they mean for hosting Which countries are in the Five, Nine and 14 Eyes, what is official and what was leaked, and what membership really means for a server in each of our seven locations.26 September 2026 8 min read](https://offshoreserv.com/blog/14-eyes-countries-hosting)

---

OffshoreServ is an offshore hosting provider: VPS, dedicated servers, Windows RDP and GPU servers in seven jurisdictions (Iceland, Switzerland, Moldova, Romania, the Netherlands, Bulgaria and Malaysia), paid only in cryptocurrency (Bitcoin, Ethereum, Monero, Tether (USDT) and Solana), with no identity checks (no KYC).

Prices and plans: https://offshoreserv.com/pricing · Answers: https://offshoreserv.com/faq · Every page: https://offshoreserv.com/llms.txt
