On this page

To
This guide takes an empty GPU server to a private endpoint: the card to rent, exact commands for Ollama and vLLM, TLS and authentication, benchmarks and cost. Commands assume
Why self-host an LLM?
Your prompts stay on your server. With a hosted API, every prompt and answer passes through someone else's infrastructure, under their logging and retention rules. A private LLM server runs the model on your own GPU. Ollama's FAQ puts it plainly: when you run locally, "we don't see your prompts or data".
The cost is predictable. A monthly server costs the same for ten requests or ten million: cheaper than
You control the model: its version, its quantization and its context length. It does not change when a provider updates its catalog, and you can
The
Pick the GPU by VRAM
To run an LLM on a GPU server, start from memory: the weights, the KV cache for your context and some headroom must all fit in VRAM. These figures from our guide on how much VRAM LLMs need include an
| Model (example) | A GPU that fits | |||
|---|---|---|---|---|
| 7B to 8B ( | About | About | About | |
| 13B to 14B ( | About | About | About | |
| 32B ( | About | About | About | |
| 70B ( | About | About | About | A100 or H100 at |
| Fits one | Not applicable | Not applicable | A100 or H100 |
An
Quick start: self-host an LLM with Ollama
Ollama's Linux install script sets up a system service and detects the NVIDIA GPU. You can read it before running it:
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
ollama run llama3.1:8b
ollama ps
llama3.1:8bollama run/byeollama psOLLAMA_CONTEXT_LENGTH
ss -tln | grep 11434
curl http://127.0.0.1:11434/api/generate -d '{"model": "llama3.1:8b", "prompt": "Why is the sky blue?", "stream": false}'
Ollama binds 127.0.0.1:11434OLLAMA_HOST=0.0.0.0
Ollama also serves /v1http://127.0.0.1:11434/v1sudo systemctl edit ollama.service
[Service]
Environment="OLLAMA_NO_CLOUD=1"
Environment="OLLAMA_CONTEXT_LENGTH=16384"
sudo systemctl daemon-reload
sudo systemctl restart ollama
Serve an OpenAI-compatible API with vLLM
vLLM is built for throughput: it batches many requests on the GPU at once and speaks the OpenAI protocol, so existing client code works unchanged. Install it in its own virtual environment as the vLLM installation guide shows, then create a key and start the server:
python3 -m venv ~/venvs/vllm
source ~/venvs/vllm/bin/activate
pip install --upgrade pip
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
The key is required because VLLM_API_KEY--api-keyhttp://localhost:8000--host--max-model-lennvidia-smi
from openai import OpenAI
client = OpenAI(base_url="http://127.0.0.1:8000/v1", api_key="YOUR_KEY")
reply = client.chat.completions.create(
model="Qwen/Qwen2.5-7B-Instruct",
messages=[{"role": "user", "content": "Explain RAID 1 in one sentence."}],
)
print(reply.choices[0].message.content)
The same code talks to Ollama at http://127.0.0.1:11434/v1--ipc=host
sudo docker run --rm --gpus all --ipc=host -p 127.0.0.1:8000:8000 \
-v ~/.cache/huggingface:/root/.cache/huggingface \
vllm/vllm-openai:latest --model Qwen/Qwen2.5-7B-Instruct --api-key "$VLLM_API_KEY"
On two or four GPUs, --tensor-parallel-size 24
Ollama vs vLLM at a glance
| Question | Ollama | vLLM |
|---|---|---|
| Setup | One script, runs as a system service | Python package or Docker image |
| Models | Ollama library, by name | Hugging Face, by repository |
| Default address | 127.0.0.1:11434 | --host |
| Authentication | None on the local API | Optional key that covers only some routes |
| Built for | One user or a small team, quick tests | Many concurrent requests, |
Put TLS and authentication in front
Neither server should face the internet as it is. Ollama's documentation states that the local API does not require authentication. vLLM's security guide says --api-key/v1/v2/inference/cohere/pause/tokenize
For yourself: an SSH tunnel
ssh -N -L 11434:127.0.0.1:11434 -L 8000:127.0.0.1:8000 YOUR_USER@SERVER_IP
While it runs, clients on your computer use http://127.0.0.1:11434/v1http://127.0.0.1:8000/v1
For apps and teams: Caddy with TLS and a key
Caddy gets and renews a certificate by itself and redirects HTTP to HTTPS. Point a domain's A record at the server, open ports 80 and 443 in the firewall, install Caddy from its official repository, make a key with openssl rand -hex 32/etc/caddy/Caddyfile
llm.example.com {
@denied {
not {
path /v1/*
header Authorization "Bearer YOUR_LONG_RANDOM_KEY"
}
}
respond @denied 403
reverse_proxy 127.0.0.1:11434 {
header_up Host localhost:11434
}
}
Then run sudo systemctl reload caddy/v1/respondreverse_proxy/api127.0.0.1:8000header_up
Clients then use https://llm.example.com/v1
Measure throughput yourself
Tokens per second depend on the card, model, quantization, engine version, context length and concurrency, so we publish no figures we have not measured. For one user on a dense model, the ceiling is memory bandwidth divided by the size of the weights: about
With Ollama
ollama run llama3.1:8b --verbose
--verboseeval_counteval_durationeval_count / eval_duration × 10^9
With vLLM
vLLM ships a load generator. This sends OPENAI_API_KEY
export OPENAI_API_KEY="$VLLM_API_KEY"
vllm bench serve --model Qwen/Qwen2.5-7B-Instruct --dataset-name random --random-input-len 1024 --random-output-len 256 --num-prompts 200 --max-concurrency 16
Read three lines of the report: output token throughput for the whole server, mean TTFT (time to first token) for responsiveness and mean TPOT (time per output token) for what each user sees. Repeat at concurrency 1, 8 and 32, watch nvidia-smi dmon -s pucm
Keep prompts private
- No public ports.
must show 11434 and 8000 on 127.0.0.1 only. Allow SSH, plus 80 and 443 for Caddy, and nothing else, as in our hardening guide; publish Docker ports assudo ss -tlnp .127.0.0.1:PORT:PORT - No cloud models by accident. Ollama can also run cloud models, such as
, on its own servers, which then process your prompts.gemma4:cloud disables cloud models and web search.OLLAMA_NO_CLOUD=1 - Telemetry off. vLLM collects anonymous usage statistics on your hardware and configuration by default;
orVLLM_NO_USAGE_STATS=1 stops them.DO_NOT_TRACK=1 turns off Hugging Face telemetry, and once the model is downloaded,HF_HUB_DISABLE_TELEMETRY=1 stops all calls to the Hub.HF_HUB_OFFLINE=1 - Quiet logs. vLLM logs requests only with
, and prompt text only at DEBUG level. Keep it that way, and check what any chat front end stores.--enable-log-requests
export VLLM_NO_USAGE_STATS=1
export HF_HUB_DISABLE_TELEMETRY=1
export HF_HUB_OFFLINE=1
For a service, put these in its environment file. On our side, we do not log or inspect the traffic of customer servers: no deep packet inspection, no content scanning, and we never look at what runs inside them.
Cost: API vs cloud GPU vs your own server
- An API bills per token: nothing to run and no idle cost, but every prompt goes to the provider and the bill grows with volume.
- A cloud GPU bills by the hour. Against a monthly server, the
break-even is the monthly price divided by the hourly rate: our $84.99RTX 4090 equals85 hours at an example $1.00 an hour. - Your own server costs the same busy or idle, so the cost per token falls as use grows. Around the clock, our
RTX 4090 works out to about $0.116 an hour and the H100 to about $0.797.
Against an API, the server pays for itself above this average throughput:
break-even tokens per second = monthly price ÷ (API price per million tokens × 2.628)
2.628 is the number of millions of seconds in a
Quarterly billing saves 5% and yearly billing 12%, which brings the
Frequently asked questions
Can I run an LLM on my own server?
Yes. Any Linux server whose NVIDIA GPU has enough VRAM for the model works: about
Is Ollama or vLLM better?
They solve different problems. Ollama is quicker to set up and suits one user or a small team. vLLM batches many concurrent requests, splits models across GPUs and checks an API key, which suits applications. Both answer
How much VRAM do I need for a 70B model?
About
Does Ollama have authentication?
Not on its local API. Ollama's documentation states that the API at localhost:11434 does not require authentication; API keys apply only to its cloud service. Keep the default binding to 127.0.0.1, which only the server itself can reach, and add an SSH tunnel or a reverse proxy with TLS and a key check before anyone else connects.
Is self-hosting an LLM cheaper than an API?
At steady volume, often yes; at low or irregular volume, usually not. Divide the server's monthly price by your API's price per million tokens times 2.628: the result is the average tokens per second it must sustain to break even. Privacy can matter more than price: on your own server, prompts stay under your control.
Offshore VPS, dedicated, RDP and GPU servers in seven jurisdictions.


