Launch pricing: every plan costs 30% less than the cheapest offshore competitor we track. See the benchmarkEvery plan 30% under the cheapest offshore host

AI & GPUs

How to self-host an LLM on a GPU server (Ollama, vLLM)

From an empty GPU server to a private, OpenAI-compatible endpoint: which card to rent, Ollama and vLLM setup, TLS and keys, benchmarks, privacy and cost.

10 min readBy the OffshoreServ team

Key takeaways

  • Size by VRAM: weights, plus the cache for your context, plus 10 to 20% headroom. A 4-bit 70B model needs about 50 GB.
  • Ollama runs a model in two commands; its API on 127.0.0.1:11434 has no authentication.
  • vLLM batches many users behind an OpenAI-compatible API on port 8000; its optional key covers only some routes.
  • Expose nothing directly: use an SSH tunnel, or a TLS reverse proxy that checks a key and allows only /v1 paths.
  • Benchmark on your own server, switch off telemetry, and compare the monthly price with your API bill.
On this page
  1. Why self-host an LLM?
  2. Pick the GPU by VRAM
  3. Quick start: self-host an LLM with Ollama
  4. Serve an OpenAI-compatible API with vLLM
  5. Put TLS and authentication in front
  6. Measure throughput yourself
  7. Keep prompts private
  8. Cost: API vs cloud GPU vs your own server
  9. Frequently asked questions

To self-host an LLM, rent a GPU server with enough VRAM for the model, install Ollama for a quick start or vLLM for an OpenAI-compatible API, and keep both bound to 127.0.0.1. Reach them through an SSH tunnel or a TLS reverse proxy that checks a key. An 8B model runs at full precision on a 24 GB RTX 4090.

This guide takes an empty GPU server to a private endpoint: the card to rent, exact commands for Ollama and vLLM, TLS and authentication, benchmarks and cost. Commands assume Ubuntu 24.04 with the NVIDIA driver and CUDA installed, as on our GPU servers.

Why self-host an LLM?

Your prompts stay on your server. With a hosted API, every prompt and answer passes through someone else's infrastructure, under their logging and retention rules. A private LLM server runs the model on your own GPU. Ollama's FAQ puts it plainly: when you run locally, "we don't see your prompts or data".

The cost is predictable. A monthly server costs the same for ten requests or ten million: cheaper than per-token pricing at steady use, more expensive at low use.

You control the model: its version, its quantization and its context length. It does not change when a provider updates its catalog, and you can fine-tune it on data that stays on the server.

The trade-offs: you handle updates, security and monitoring, and open-weight models are not the proprietary models behind the largest APIs. If you need one of those, its API is the only way in.

Pick the GPU by VRAM

To run an LLM on a GPU server, start from memory: the weights, the KV cache for your context and some headroom must all fit in VRAM. These figures from our guide on how much VRAM LLMs need include an 8K-token context for one user and 10% headroom.

Model (example)4-bit8-bit16-bitA GPU that fits
7B to 8B (Llama 3.1 8B)About 7 GBAbout 11 GBAbout 19 GBRTX 4090, even at 16-bit
13B to 14B (Qwen3-14B)About 11 GBAbout 18 GBAbout 32 GBRTX 4090 at 8-bit, RTX 5090 for long contexts
32B (Qwen3-32B)About 24 GBAbout 40 GBAbout 73 GBRTX 5090 at 4-bit
70B (Llama 3.3 70B)About 50 GBAbout 85 GBAbout 157 GBA100 or H100 at 4-bit, 2 × H100 at 8-bit
gpt-oss-120b (mixture of experts)Fits one 80 GB GPU (native MXFP4)Not applicableNot applicableA100 or H100

An RTX 4090 server costs $84.99 a month, an RTX 5090 server $134.99, the A100 80 GB $355.99 and an H100 server $581.99. A 4-bit 32B model needs about 24 GB, which leaves no room on a 24 GB card once the context grows, so take the RTX 5090 for it. Between the two 80 GB cards, the H100's 3.35 TB/s of memory bandwidth makes it generate faster than the A100 on the same model.

Quick start: self-host an LLM with Ollama

Ollama's Linux install script sets up a system service and detects the NVIDIA GPU. You can read it before running it:

curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
ollama run llama3.1:8b
ollama ps

llama3.1:8b is a 4.9 GB download. ollama run opens a chat (/bye leaves it), and ollama ps should show 100% GPU; a CPU share means the model does not fit in VRAM. Its CONTEXT column shows the context length, which Ollama sizes from the available VRAM (4K tokens below 24 GiB, 32K from 24 GiB, 256K from 48 GiB) and OLLAMA_CONTEXT_LENGTH overrides. Then check where it listens and call its API:

ss -tln | grep 11434
curl http://127.0.0.1:11434/api/generate -d '{"model": "llama3.1:8b", "prompt": "Why is the sky blue?", "stream": false}'

Ollama binds 127.0.0.1 port 11434 by default, so the first line should show 127.0.0.1:11434. Leave it there. The FAQ explains how to expose it with OLLAMA_HOST=0.0.0.0, but the local API has no authentication: anyone who reached the port could run, pull and delete models.

Ollama also serves OpenAI-style routes under /v1. Clients need the base URL http://127.0.0.1:11434/v1 and any API key, which Ollama ignores. Settings are environment variables of the service: run sudo systemctl edit ollama.service, add these lines, then reload and restart:

[Service]
Environment="OLLAMA_NO_CLOUD=1"
Environment="OLLAMA_CONTEXT_LENGTH=16384"
sudo systemctl daemon-reload
sudo systemctl restart ollama

Serve an OpenAI-compatible API with vLLM

vLLM is built for throughput: it batches many requests on the GPU at once and speaks the OpenAI protocol, so existing client code works unchanged. Install it in its own virtual environment as the vLLM installation guide shows, then create a key and start the server:

python3 -m venv ~/venvs/vllm
source ~/venvs/vllm/bin/activate
pip install --upgrade pip
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192

The key is required because VLLM_API_KEY is set; the --api-key flag does the same. The quickstart gives http://localhost:8000 as the default address, but the server code binds all IPv4 interfaces when no host is given, so always pass --host. --max-model-len caps the context and its cache. vLLM reserves 92% of GPU memory by default, so a nearly full card in nvidia-smi is normal. The model downloads from Hugging Face on first start and needs about 15 GB for its weights: use a 24 GB card or larger. Test it with the official OpenAI client:

from openai import OpenAI

client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="YOUR_KEY")
reply = client.chat.completions.create(
    model="Qwen/Qwen2.5-7B-Instruct",
    messages=[{"role": "user", "content": "Explain RAID 1 in one sentence."}],
)
print(reply.choices[0].message.content)

The same code talks to Ollama at http://127.0.0.1:11434/v1. With Docker, publish the port on 127.0.0.1 only, since Docker's port rules bypass ufw, and keep --ipc=host for the shared memory PyTorch uses:

sudo docker run --rm --gpus all --ipc=host -p 127.0.0.1:8000:8000 \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  vllm/vllm-openai:latest --model Qwen/Qwen2.5-7B-Instruct --api-key "$VLLM_API_KEY"

On two or four GPUs, --tensor-parallel-size 2 or 4 splits one model across them. Our GPU images guide adds the NVIDIA Container Toolkit and a systemd unit that keeps vLLM running.

Ollama vs vLLM at a glance

QuestionOllamavLLM
SetupOne script, runs as a system servicePython package or Docker image
ModelsOllama library, by nameHugging Face, by repository
Default address127.0.0.1:11434Port 8000 on all interfaces, unless you set --host
AuthenticationNone on the local APIOptional key that covers only some routes
Built forOne user or a small team, quick testsMany concurrent requests, multi-GPU serving

Put TLS and authentication in front

Neither server should face the internet as it is. Ollama's documentation states that the local API does not require authentication. vLLM's security guide says --api-key protects only routes under /v1, /v2, /inference and /cohere, leaves others such as /pause and /tokenize open, and warns against relying on the key alone.

For yourself: an SSH tunnel

ssh -N -L 11434:127.0.0.1:11434 -L 8000:127.0.0.1:8000 YOUR_USER@SERVER_IP

While it runs, clients on your computer use http://127.0.0.1:11434/v1 or http://127.0.0.1:8000/v1. SSH encrypts the traffic and your key authenticates you; the server opens no new port. The command also works in PowerShell.

For apps and teams: Caddy with TLS and a key

Caddy gets and renews a certificate by itself and redirects HTTP to HTTPS. Point a domain's A record at the server, open ports 80 and 443 in the firewall, install Caddy from its official repository, make a key with openssl rand -hex 32, and replace /etc/caddy/Caddyfile with:

llm.example.com {
	@denied {
		not {
			path /v1/*
			header Authorization "Bearer YOUR_LONG_RANDOM_KEY"
		}
	}
	respond @denied 403
	reverse_proxy 127.0.0.1:11434 {
		header_up Host localhost:11434
	}
}

Then run sudo systemctl reload caddy. Requests that are not both under /v1/ and carrying the key get a 403, because respond runs before reverse_proxy. The path rule also keeps Ollama's native /api routes, which can pull and delete models, private. The Host rewrite matters: listening on 127.0.0.1, Ollama answers 403 to requests for other host names, and its FAQ's nginx example rewrites the header the same way. For vLLM, proxy to 127.0.0.1:8000 without the header_up block and use the server's key.

Clients then use https://llm.example.com/v1 and the key, which OpenAI-compatible clients send as a Bearer token; Caddy flushes streamed answers immediately. A service for other people also needs rate limits, and our acceptable use policy applies to what it generates.

Measure throughput yourself

Tokens per second depend on the card, model, quantization, engine version, context length and concurrency, so we publish no figures we have not measured. For one user on a dense model, the ceiling is memory bandwidth divided by the size of the weights: about 205 tokens per second for a 4-bit 8B model on an RTX 4090, per our VRAM guide. Real results land below it.

With Ollama

ollama run llama3.1:8b --verbose

--verbose prints timings after each answer, including the eval rate, the generation speed in tokens per second. Through the API, responses carry eval_count and eval_duration in nanoseconds, and Ollama's API reference gives the rate as eval_count / eval_duration × 10^9.

With vLLM

vLLM ships a load generator. This sends 200 requests of 1,024 random input tokens and 256 output tokens, 16 at a time, to 127.0.0.1:8000, with OPENAI_API_KEY as the Bearer token:

export OPENAI_API_KEY="$VLLM_API_KEY"
vllm bench serve --model Qwen/Qwen2.5-7B-Instruct --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 200 --max-concurrency 16

Read three lines of the report: output token throughput for the whole server, mean TTFT (time to first token) for responsiveness and mean TPOT (time per output token) for what each user sees. Repeat at concurrency 1, 8 and 32, watch nvidia-smi dmon -s pucm meanwhile, and use prompt lengths close to your real traffic.

Keep prompts private

Self-hosting keeps prompts on your server only if nothing else sends them away:

  1. No public ports. sudo ss -tlnp must show 11434 and 8000 on 127.0.0.1 only. Allow SSH, plus 80 and 443 for Caddy, and nothing else, as in our hardening guide; publish Docker ports as 127.0.0.1:PORT:PORT.
  2. No cloud models by accident. Ollama can also run cloud models, such as gemma4:cloud, on its own servers, which then process your prompts. OLLAMA_NO_CLOUD=1 disables cloud models and web search.
  3. Telemetry off. vLLM collects anonymous usage statistics on your hardware and configuration by default; VLLM_NO_USAGE_STATS=1 or DO_NOT_TRACK=1 stops them. HF_HUB_DISABLE_TELEMETRY=1 turns off Hugging Face telemetry, and once the model is downloaded, HF_HUB_OFFLINE=1 stops all calls to the Hub.
  4. Quiet logs. vLLM logs requests only with --enable-log-requests, and prompt text only at DEBUG level. Keep it that way, and check what any chat front end stores.
export VLLM_NO_USAGE_STATS=1
export HF_HUB_DISABLE_TELEMETRY=1
export HF_HUB_OFFLINE=1

For a service, put these in its environment file. On our side, we do not log or inspect the traffic of customer servers: no deep packet inspection, no content scanning, and we never look at what runs inside them.

Cost: API vs cloud GPU vs your own server

  • An API bills per token: nothing to run and no idle cost, but every prompt goes to the provider and the bill grows with volume.
  • A cloud GPU bills by the hour. Against a monthly server, the break-even is the monthly price divided by the hourly rate: our $84.99 RTX 4090 equals 85 hours at an example $1.00 an hour.
  • Your own server costs the same busy or idle, so the cost per token falls as use grows. Around the clock, our RTX 4090 works out to about $0.116 an hour and the H100 to about $0.797.

Against an API, the server pays for itself above this average throughput:

break-even tokens per second = monthly price ÷ (API price per million tokens × 2.628)

2.628 is the number of millions of seconds in a 730-hour month. At an example API price of $1 per million tokens, an $84.99 server breaks even at an average of about 32 tokens per second, idle time included, and a $581.99 H100 at about 221. These are example prices, not quotes, and APIs bill input tokens too. Our comparison of dedicated GPU servers and cloud GPUs covers the hourly side.

Quarterly billing saves 5% and yearly billing 12%, which brings the RTX 4090 to $74.79 a month. Servers are ready in 1 to 24 hours, and GPU servers are not refundable once delivered, so size the model before you order.

Frequently asked questions

Can I run an LLM on my own server?

Yes. Any Linux server whose NVIDIA GPU has enough VRAM for the model works: about 7 GB for an 8B model and about 50 GB for a 70B model, both at 4-bit with context. Install Ollama or vLLM, keep it on 127.0.0.1, and connect through an SSH tunnel or a TLS reverse proxy that checks a key.

Is Ollama or vLLM better?

They solve different problems. Ollama is quicker to set up and suits one user or a small team. vLLM batches many concurrent requests, splits models across GPUs and checks an API key, which suits applications. Both answer OpenAI-style requests under /v1, so you can start on Ollama and move to vLLM without rewriting your client.

How much VRAM do I need for a 70B model?

About 43 GB for the weights at 4-bit, or about 50 GB with an 8K context and headroom: one 80 GB A100 or H100, or two 32 GB RTX 5090s. At 8-bit, plan for about 85 GB, which means two H100s; at 16-bit, about 157 GB, on a multi-GPU H100 server. Longer contexts and more users need more.

Does Ollama have authentication?

Not on its local API. Ollama's documentation states that the API at localhost:11434 does not require authentication; API keys apply only to its cloud service. Keep the default binding to 127.0.0.1, which only the server itself can reach, and add an SSH tunnel or a reverse proxy with TLS and a key check before anyone else connects.

Is self-hosting an LLM cheaper than an API?

At steady volume, often yes; at low or irregular volume, usually not. Divide the server's monthly price by your API's price per million tokens times 2.628: the result is the average tokens per second it must sustain to break even. Privacy can matter more than price: on your own server, prompts stay under your control.

Host it where the law is on your side.

Offshore VPS, dedicated, RDP and GPU servers in seven jurisdictions. No KYC, paid in crypto.

Welcome back

Sign in to manage your servers and your balance.

No KYCHuman check by Cloudflare TurnstileNo tracking