All guides
Getting started
VPS
- VPS operating systems
- VPS snapshots and
off-site backups - How to host your own VPN on a VPS with WireGuard
- How to run a Tor relay, bridge or onion service on a VPS
- How to
self-host BTCPay Server on a VPS - How to run a Bitcoin or Monero node on a VPS
Dedicated servers
- Using IPMI and the KVM console on a dedicated server
- Choosing a RAID layout for your dedicated server
Windows RDP
Windows Server 2019 , 2022 or 2025 vsWindows 10 and 11 for RDP- How to connect to a Windows RDP server from any device
GPU servers
Security
On this page
Our GPU servers come with dedicated NVIDIA GPUs and a Linux image that is ready for CUDA work. This guide checks the image, then sets up the common AI tools on it. The commands assume sudoSERVER_IP
What the image includes
Ubuntu 24.04 LTS , with standard support untilMay 2029 .- The NVIDIA driver, including the
tool.nvidia-smi - CUDA, including the
compiler.nvcc
Versions change as we refresh the image, so check yours:
nvidia-smi
nvcc --version
The top of the nvidia-sminvcc/usr/local/cuda
echo 'export PATH=/usr/local/cuda/bin:$PATH' >> ~/.bashrc
source ~/.bashrc
Verify the GPU
nvidia-smi -L
nvidia-smi topo -m
The first command lists every GPU with its model; check that the count and model match your plan, and that the memory shown by nvidia-smiCUDA_VISIBLE_DEVICES
CUDA_VISIBLE_DEVICES=1 python train.py
Prepare Python
pip installexternally-managed-environment
sudo apt update
sudo apt install -y python3-venv python3-dev build-essential git
Install PyTorch
PyTorch publishes builds for specific CUDA versions. Current releases offer nvidia-smi
python3 -m venv ~/venvs/torch
source ~/venvs/torch/bin/activate
pip install --upgrade pip
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"
The last line should print Truecu130cu126
Run models with Ollama
For the why and the security side of a private model server, read how to
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
ollama run llama3.1:8b
ollama ps
ollama run/byeollama ps
Ollama listens on 127.0.0.1:11434/v1/usr/share/ollama/.ollama/modelsEnvironment=
sudo systemctl edit ollama.service
Serve an OpenAI-compatible endpoint with vLLM
vLLM is an inference server built for throughput: it batches many requests at once and speaks the same HTTP protocol as OpenAI, so existing client libraries work with it. Its builds are compiled against specific PyTorch and CUDA versions, so install it in its own environment with the index URL from the vLLM documentation:
python3 -m venv ~/venvs/vllm
source ~/venvs/vllm/bin/activate
pip install --upgrade pip
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
If your driver reports an older CUDA version, the vLLM documentation also describes an install through uv--torch-backend=auto
Create a token so that only you can use the server, note it, and start the server on localhost:
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
The first start downloads the model from Hugging Face into ~/.cache/huggingfaceQwen/Qwen2.5-3B-Instruct--tensor-parallel-size 2--hostHF_TOKEN
Test it from a second SSH session, after exporting the same token there:
curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $VLLM_API_KEY" -d '{"model": "Qwen/Qwen2.5-7B-Instruct", "messages": [{"role": "user", "content": "Say hello in five words."}]}'
Run vLLM as a service
Started from a shell, the server stops when you log out. To keep it running and bring it back after a reboot, store the token in a file only root can read, then create a systemd unit:
sudo install -m 600 /dev/null /etc/vllm.env
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
sudo cat /etc/vllm.env
sudo systemctl edit --force --full vllm.service
Paste this unit into the editor, replacing YOUR_USER
[Unit]
Description=vLLM server
After=network-online.target
Wants=network-online.target
[Service]
User=YOUR_USER
EnvironmentFile=/etc/vllm.env
ExecStart=/home/YOUR_USER/venvs/vllm/bin/vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure
[Install]
WantedBy=multi-user.target
Then start it, enable it at boot and follow its log:
sudo systemctl enable --now vllm.service
sudo journalctl -u vllm.service -f
Generate images with ComfyUI
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python main.py --listen 127.0.0.1 --port 8188
Put model checkpoints in ComfyUI/models/checkpointsmodels--listen
Docker and the NVIDIA Container Toolkit
Containers need NVIDIA's Container Toolkit to reach the GPUs. Install Docker from Ubuntu's archive, add NVIDIA's repository, and configure Docker to use the NVIDIA runtime:
sudo apt install -y docker.io
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
sudo docker run --rm --gpus all ubuntu nvidia-smi
The last command should print the same table as on the host. Docker CE from Docker's own repository works the same way if you prefer it. As an example, this runs vLLM in a container and publishes it on localhost only:
sudo docker run --rm --gpus all --ipc=host -p 127.0.0.1:8000:8000 -v ~/.cache/huggingface:/root/.cache/huggingface vllm/vllm-openai:latest --model Qwen/Qwen2.5-7B-Instruct
Keep endpoints private
None of these tools is safe to expose as it is: Ollama and ComfyUI have no login at all, and vLLM's token is optional. Keep them bound to 127.0.0.1
ssh -N -L 8188:127.0.0.1:8188 -L 8000:127.0.0.1:8000 -L 11434:127.0.0.1:11434 YOUR_USER@SERVER_IP
While the tunnel runs, open http://127.0.0.1:8188http://127.0.0.1:8000/v1
Monitor the GPUs
nvidia-smi dmon -s pucm
watch -n 1 nvidia-smi
nvidia-smi --query-gpu=index,temperature.gpu,utilization.gpu,memory.used,memory.total,power.draw --format=csv -l 5
sudo apt install -y nvtop
nvtop
nvidia-smi dmonnvtop
nvidia-smi -q -d PERFORMANCE
Troubleshooting
- "CUDA out of memory": the model, its context or the batch does not fit. Choose a smaller or more compressed model, lower
in vLLM, or reduce the batch size.--max-model-len shows whether another process still holds memory.nvidia-smi returns False: the virtual environment is not active, or the PyTorch build expects a newer CUDA version than the driver supports. Comparetorch.cuda.is_available() with the CUDA version intorch.version.cuda .nvidia-smi- "Driver/library version mismatch": the driver packages were updated but the old kernel module is still loaded. Reboot.
- A container sees no GPU: add
to--gpus all , and check thatdocker run ran and Docker was restarted.nvidia-ctk runtime configure - Slow first start: the first run downloads models and compiles GPU kernels. Later starts are faster.
Ready to pick hardware? Compare the RTX servers, the data center GPUs and the
Dedicated and GPU customers can open a ticket from the client area with the server’s IP address and what they tried. First reply target: under