Launch pricing: every plan costs 30% less than the cheapest offshore competitor we track. See the benchmarkEvery plan 30% under the cheapest offshore host

Knowledge base · GPU servers

GPU server images: CUDA, PyTorch, Ollama, vLLM and ComfyUI

What the OffshoreServ GPU image includes, and how to run PyTorch, Ollama, vLLM, ComfyUI and Docker on Ubuntu 24.04, keep endpoints private and watch the GPUs.

Updated 8 min read

In this guide

  • The Linux image is Ubuntu 24.04 LTS with the NVIDIA driver and CUDA pre-installed; check both with nvidia-smi and nvcc.
  • Install Python tools in virtual environments, because Ubuntu 24.04 blocks system-wide pip installs.
  • Ollama, vLLM and ComfyUI take minutes to set up. Keep them on 127.0.0.1 and reach them through an SSH tunnel.
  • Docker needs the NVIDIA Container Toolkit, and published ports bypass ufw unless you bind them to 127.0.0.1.
GPU servers11 sections
All guides
On this page
  1. What the image includes
  2. Verify the GPU
  3. Prepare Python
  4. Install PyTorch
  5. Run models with Ollama
  6. Serve an OpenAI-compatible endpoint with vLLM
  7. Generate images with ComfyUI
  8. Docker and the NVIDIA Container Toolkit
  9. Keep endpoints private
  10. Monitor the GPUs
  11. Troubleshooting

Our GPU servers come with dedicated NVIDIA GPUs and a Linux image that is ready for CUDA work. This guide checks the image, then sets up the common AI tools on it. The commands assume Ubuntu 24.04 and a user with sudo rights; if you work as root, leave out sudo. Replace SERVER_IP with your server's address.

What the image includes

  • Ubuntu 24.04 LTS, with standard support until May 2029.
  • The NVIDIA driver, including the nvidia-smi tool.
  • CUDA, including the nvcc compiler.

Versions change as we refresh the image, so check yours:

nvidia-smi
nvcc --version

The top of the nvidia-smi output shows the driver version and a CUDA Version. That is the newest CUDA release the driver supports, which matters when you choose PyTorch builds below. If nvcc is not found, the toolkit is usually installed under /usr/local/cuda but missing from your PATH:

echo 'export PATH=/usr/local/cuda/bin:$PATH' >> ~/.bashrc
source ~/.bashrc

Verify the GPU

nvidia-smi -L
nvidia-smi topo -m

The first command lists every GPU with its model; check that the count and model match your plan, and that the memory shown by nvidia-smi matches the card, for example 24 GB on an RTX 4090. On multi-GPU servers, the second command shows how the GPUs connect to each other (NVLink or PCIe), which affects how fast they share work. To run a job on one specific GPU, set CUDA_VISIBLE_DEVICES; the process then sees only that card, numbered from 0:

CUDA_VISIBLE_DEVICES=1 python train.py

Prepare Python

Ubuntu 24.04 marks its system Python as externally managed, so pip install outside a virtual environment fails with an externally-managed-environment error. Create one environment per tool: vLLM, for example, pins its own PyTorch version. Install the basics first; the compiler and headers are needed by tools that compile GPU kernels at run time:

sudo apt update
sudo apt install -y python3-venv python3-dev build-essential git

Install PyTorch

PyTorch publishes builds for specific CUDA versions. Current releases offer CUDA 12.6 and CUDA 13.0 builds, plus an experimental CUDA 13.2 build. Choose one that is no newer than the CUDA version nvidia-smi reports:

python3 -m venv ~/venvs/torch
source ~/venvs/torch/bin/activate
pip install --upgrade pip
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"

The last line should print True and your GPU's name. The cu130 build needs a driver that reports CUDA 13.0 or newer. If yours reports 12.x, use cu126 in the index URL instead. That build covers every GPU in our range except the RTX 5090, which needs CUDA 12.8 or newer.

Run models with Ollama

For the why and the security side of a private model server, read how to self-host an LLM on a GPU server. Ollama is the quickest way to run open-weight language models. Its install script sets up a system service and detects the NVIDIA GPU. You can download the script and read it before you run it, or pipe it straight to the shell:

curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
ollama run llama3.1:8b
ollama ps

ollama run opens a chat prompt; type /bye to leave it. In the output of ollama ps, the Processor column should read 100% GPU. A CPU share means the model does not fit in video memory: choose a smaller model or a more compressed quantization. How much VRAM do you need for LLMs helps you size it.

Ollama listens on 127.0.0.1:11434 by default. Keep it that way and connect through the SSH tunnel described below. It also accepts OpenAI-style requests under /v1, so the same client code works with Ollama and vLLM. Models are stored in /usr/share/ollama/.ollama/models. To change settings such as the listen address, add Environment= lines to the service with:

sudo systemctl edit ollama.service

Serve an OpenAI-compatible endpoint with vLLM

vLLM is an inference server built for throughput: it batches many requests at once and speaks the same HTTP protocol as OpenAI, so existing client libraries work with it. Its builds are compiled against specific PyTorch and CUDA versions, so install it in its own environment with the index URL from the vLLM documentation:

python3 -m venv ~/venvs/vllm
source ~/venvs/vllm/bin/activate
pip install --upgrade pip
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129

If your driver reports an older CUDA version, the vLLM documentation also describes an install through uv with --torch-backend=auto, which picks a PyTorch build that matches the driver.

Create a token so that only you can use the server, note it, and start the server on localhost:

export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192

The first start downloads the model from Hugging Face into ~/.cache/huggingface; this model needs about 15 GB of video memory for its weights alone, so use a GPU with 24 GB or more. On the 16 GB RTX A4000, choose a smaller model such as Qwen/Qwen2.5-3B-Instruct. On servers with several GPUs, add --tensor-parallel-size 2 (or 4) to split a model across them. vLLM listens on all interfaces unless you pass --host, so always set it. Gated models, such as Meta's Llama family, also need a Hugging Face access token: accept the model's license on its Hugging Face page, then export HF_TOKEN before you start the server.

Test it from a second SSH session, after exporting the same token there:

curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $VLLM_API_KEY" -d '{"model": "Qwen/Qwen2.5-7B-Instruct", "messages": [{"role": "user", "content": "Say hello in five words."}]}'

Run vLLM as a service

Started from a shell, the server stops when you log out. To keep it running and bring it back after a reboot, store the token in a file only root can read, then create a systemd unit:

sudo install -m 600 /dev/null /etc/vllm.env
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
sudo cat /etc/vllm.env
sudo systemctl edit --force --full vllm.service

Paste this unit into the editor, replacing YOUR_USER with your username:

[Unit]
Description=vLLM server
After=network-online.target
Wants=network-online.target

[Service]
User=YOUR_USER
EnvironmentFile=/etc/vllm.env
ExecStart=/home/YOUR_USER/venvs/vllm/bin/vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure

[Install]
WantedBy=multi-user.target

Then start it, enable it at boot and follow its log:

sudo systemctl enable --now vllm.service
sudo journalctl -u vllm.service -f

Generate images with ComfyUI

git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python main.py --listen 127.0.0.1 --port 8188

Put model checkpoints in ComfyUI/models/checkpoints, and other files such as VAEs and LoRAs in their matching folders under models. With --listen and no address, ComfyUI listens on all IPv4 and IPv6 interfaces, without any password; only do that behind a firewall rule that allows your IP alone. Custom nodes run arbitrary Python code with your user's rights, so install only nodes from sources you trust.

Docker and the NVIDIA Container Toolkit

Containers need NVIDIA's Container Toolkit to reach the GPUs. Install Docker from Ubuntu's archive, add NVIDIA's repository, and configure Docker to use the NVIDIA runtime:

sudo apt install -y docker.io
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
sudo docker run --rm --gpus all ubuntu nvidia-smi

The last command should print the same table as on the host. Docker CE from Docker's own repository works the same way if you prefer it. As an example, this runs vLLM in a container and publishes it on localhost only:

sudo docker run --rm --gpus all --ipc=host -p 127.0.0.1:8000:8000 -v ~/.cache/huggingface:/root/.cache/huggingface vllm/vllm-openai:latest --model Qwen/Qwen2.5-7B-Instruct

Keep endpoints private

None of these tools is safe to expose as it is: Ollama and ComfyUI have no login at all, and vLLM's token is optional. Keep them bound to 127.0.0.1 and reach them through an SSH tunnel from your own computer (on Windows, the same command works in PowerShell):

ssh -N -L 8188:127.0.0.1:8188 -L 8000:127.0.0.1:8000 -L 11434:127.0.0.1:11434 YOUR_USER@SERVER_IP

While the tunnel runs, open http://127.0.0.1:8188 in your browser for ComfyUI, and point OpenAI-compatible clients at http://127.0.0.1:8000/v1. Close the rest of the server with a firewall that allows only SSH, as in the hardening guide. If a service must be public, put a reverse proxy with TLS, authentication and rate limits in front of it, and keep the acceptable use policy in mind.

Monitor the GPUs

nvidia-smi dmon -s pucm
watch -n 1 nvidia-smi
nvidia-smi --query-gpu=index,temperature.gpu,utilization.gpu,memory.used,memory.total,power.draw --format=csv -l 5
sudo apt install -y nvtop
nvtop

nvidia-smi dmon prints one line per GPU every second, with power and temperature (p), utilization (u), clocks (c) and memory (m). The query form writes CSV every five seconds, handy for logging long jobs. nvtop is an interactive view of all GPUs and their processes, similar to htop. If clocks drop under load, check why:

nvidia-smi -q -d PERFORMANCE

Troubleshooting

  • "CUDA out of memory": the model, its context or the batch does not fit. Choose a smaller or more compressed model, lower --max-model-len in vLLM, or reduce the batch size. nvidia-smi shows whether another process still holds memory.
  • torch.cuda.is_available() returns False: the virtual environment is not active, or the PyTorch build expects a newer CUDA version than the driver supports. Compare torch.version.cuda with the CUDA version in nvidia-smi.
  • "Driver/library version mismatch": the driver packages were updated but the old kernel module is still loaded. Reboot.
  • A container sees no GPU: add --gpus all to docker run, and check that nvidia-ctk runtime configure ran and Docker was restarted.
  • Slow first start: the first run downloads models and compiles GPU kernels. Later starts are faster.

Ready to pick hardware? Compare the RTX servers, the data center GPUs and the multi-GPU servers.

Stuck on a step?

Dedicated and GPU customers can open a ticket from the client area with the server’s IP address and what they tried. First reply target: under 12 hours. For every other server, use the Server actions on its page, the guides and the network status page.

Welcome back

Sign in to manage your servers and your balance.

No KYCHuman check by Cloudflare TurnstileNo tracking