---
title: "GPU Server Images: CUDA, PyTorch, Ollama, vLLM, ComfyUI"
description: "What the GPU image includes, and how to run PyTorch, Ollama, vLLM, ComfyUI and Docker on Ubuntu 24.04, keep endpoints private and watch GPUs."
url: https://offshoreserv.com/docs/gpu/images
lang: en
updated: 2026-09-26
source: HTML page at the url above (canonical); this is its Markdown version
---

Knowledge base · GPU servers

# GPU server images: CUDA, PyTorch, Ollama, vLLM and ComfyUI

What the OffshoreServ GPU image includes, and how to run PyTorch, Ollama, vLLM, ComfyUI and Docker on Ubuntu 24.04, keep endpoints private and watch the GPUs.

Updated 26 September 2026 8 min read

In this guide

- The Linux image is Ubuntu 24.04 LTS with the NVIDIA driver and CUDA pre-installed; check both with nvidia-smi and nvcc.
- Install Python tools in virtual environments, because Ubuntu 24.04 blocks system-wide pip installs.
- Ollama, vLLM and ComfyUI take minutes to set up. Keep them on 127.0.0.1 and reach them through an SSH tunnel.
- Docker needs the NVIDIA Container Toolkit, and published ports bypass ufw unless you bind them to 127.0.0.1.

GPU servers11 sections

Our [GPU servers](https://offshoreserv.com/offshore-gpu-servers) come with dedicated NVIDIA GPUs and a Linux image that is ready for CUDA work. This guide checks the image, then sets up the common AI tools on it. The commands assume Ubuntu 24.04 and a user with sudo rights; if you work as root, leave out `sudo`. Replace `SERVER_IP` with your server's address.

## What the image includes

- Ubuntu 24.04 LTS, with standard support until May 2029.
- The NVIDIA driver, including the `nvidia-smi` tool.
- CUDA, including the `nvcc` compiler.

Versions change as we refresh the image, so check yours:

```
nvidia-smi
nvcc --version
```

The top of the `nvidia-smi` output shows the driver version and a **CUDA Version**. That is the newest CUDA release the driver supports, which matters when you choose PyTorch builds below. If `nvcc` is not found, the toolkit is usually installed under `/usr/local/cuda` but missing from your PATH:

```
echo 'export PATH=/usr/local/cuda/bin:$PATH' >> ~/.bashrc
source ~/.bashrc
```

> When a system update replaces the NVIDIA driver packages, `nvidia-smi` reports "Driver/library version mismatch" until you reboot. Plan driver updates together with a reboot. If you use unattended-upgrades, consider adding the NVIDIA packages to its `Package-Blacklist` and updating them by hand.

## Verify the GPU

```
nvidia-smi -L
nvidia-smi topo -m
```

The first command lists every GPU with its model; check that the count and model match your plan, and that the memory shown by `nvidia-smi` matches the card, for example 24 GB on an RTX 4090. On multi-GPU servers, the second command shows how the GPUs connect to each other (NVLink or PCIe), which affects how fast they share work. To run a job on one specific GPU, set `CUDA_VISIBLE_DEVICES`; the process then sees only that card, numbered from 0:

```
CUDA_VISIBLE_DEVICES=1 python train.py
```

## Prepare Python

Ubuntu 24.04 marks its system Python as externally managed, so `pip install` outside a virtual environment fails with an `externally-managed-environment` error. Create one environment per tool: vLLM, for example, pins its own PyTorch version. Install the basics first; the compiler and headers are needed by tools that compile GPU kernels at run time:

```
sudo apt update
sudo apt install -y python3-venv python3-dev build-essential git
```

## Install PyTorch

PyTorch publishes builds for specific CUDA versions. Current releases offer CUDA 12.6 and CUDA 13.0 builds, plus an experimental CUDA 13.2 build. Choose one that is no newer than the CUDA version `nvidia-smi` reports:

```
python3 -m venv ~/venvs/torch
source ~/venvs/torch/bin/activate
pip install --upgrade pip
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu130
python -c "import torch; print(torch.__version__, torch.cuda.is_available(), torch.cuda.get_device_name(0))"
```

The last line should print `True` and your GPU's name. The `cu130` build needs a driver that reports CUDA 13.0 or newer. If yours reports 12.x, use `cu126` in the index URL instead. That build covers every GPU in our range except the RTX 5090, which needs CUDA 12.8 or newer.

## Run models with Ollama

For the why and the security side of a private model server, read [how to self-host an LLM on a GPU server](https://offshoreserv.com/blog/self-host-llm-gpu-server). Ollama is the quickest way to run open-weight language models. Its install script sets up a system service and detects the NVIDIA GPU. You can download the script and read it before you run it, or pipe it straight to the shell:

```
curl -fsSL https://ollama.com/install.sh | sh
ollama pull llama3.1:8b
ollama run llama3.1:8b
ollama ps
```

`ollama run` opens a chat prompt; type `/bye` to leave it. In the output of `ollama ps`, the Processor column should read **100% GPU**. A CPU share means the model does not fit in video memory: choose a smaller model or a more compressed quantization. [How much VRAM do you need for LLMs](https://offshoreserv.com/blog/how-much-vram-for-llms) helps you size it.

Ollama listens on `127.0.0.1:11434` by default. Keep it that way and connect through the SSH tunnel described below. It also accepts OpenAI-style requests under `/v1`, so the same client code works with Ollama and vLLM. Models are stored in `/usr/share/ollama/.ollama/models`. To change settings such as the listen address, add `Environment=` lines to the service with:

```
sudo systemctl edit ollama.service
```

## Serve an OpenAI-compatible endpoint with vLLM

vLLM is an inference server built for throughput: it batches many requests at once and speaks the same HTTP protocol as OpenAI, so existing client libraries work with it. Its builds are compiled against specific PyTorch and CUDA versions, so install it in its own environment with the index URL from the vLLM documentation:

```
python3 -m venv ~/venvs/vllm
source ~/venvs/vllm/bin/activate
pip install --upgrade pip
pip install vllm --extra-index-url https://download.pytorch.org/whl/cu129
```

If your driver reports an older CUDA version, the vLLM documentation also describes an install through `uv` with `--torch-backend=auto`, which picks a PyTorch build that matches the driver.

Create a token so that only you can use the server, note it, and start the server on localhost:

```
export VLLM_API_KEY=$(openssl rand -hex 32)
echo "$VLLM_API_KEY"
vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
```

The first start downloads the model from Hugging Face into `~/.cache/huggingface`; this model needs about 15 GB of video memory for its weights alone, so use a GPU with 24 GB or more. On the 16 GB RTX A4000, choose a smaller model such as `Qwen/Qwen2.5-3B-Instruct`. On servers with several GPUs, add `--tensor-parallel-size 2` (or 4) to split a model across them. vLLM listens on all interfaces unless you pass `--host`, so always set it. Gated models, such as Meta's Llama family, also need a Hugging Face access token: accept the model's license on its Hugging Face page, then export `HF_TOKEN` before you start the server.

Test it from a second SSH session, after exporting the same token there:

```
curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions -H "Content-Type: application/json" -H "Authorization: Bearer $VLLM_API_KEY" -d '{"model": "Qwen/Qwen2.5-7B-Instruct", "messages": [{"role": "user", "content": "Say hello in five words."}]}'
```

### Run vLLM as a service

Started from a shell, the server stops when you log out. To keep it running and bring it back after a reboot, store the token in a file only root can read, then create a systemd unit:

```
sudo install -m 600 /dev/null /etc/vllm.env
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
sudo cat /etc/vllm.env
sudo systemctl edit --force --full vllm.service
```

Paste this unit into the editor, replacing `YOUR_USER` with your username:

```
[Unit]
Description=vLLM server
After=network-online.target
Wants=network-online.target

[Service]
User=YOUR_USER
EnvironmentFile=/etc/vllm.env
ExecStart=/home/YOUR_USER/venvs/vllm/bin/vllm serve Qwen/Qwen2.5-7B-Instruct --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure

[Install]
WantedBy=multi-user.target
```

Then start it, enable it at boot and follow its log:

```
sudo systemctl enable --now vllm.service
sudo journalctl -u vllm.service -f
```

## Generate images with ComfyUI

```
git clone https://github.com/Comfy-Org/ComfyUI.git
cd ComfyUI
python3 -m venv .venv
source .venv/bin/activate
pip install --upgrade pip
pip install torch torchvision torchaudio --extra-index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
python main.py --listen 127.0.0.1 --port 8188
```

Put model checkpoints in `ComfyUI/models/checkpoints`, and other files such as VAEs and LoRAs in their matching folders under `models`. With `--listen` and no address, ComfyUI listens on all IPv4 and IPv6 interfaces, without any password; only do that behind a firewall rule that allows your IP alone. Custom nodes run arbitrary Python code with your user's rights, so install only nodes from sources you trust.

## Docker and the NVIDIA Container Toolkit

Containers need NVIDIA's Container Toolkit to reach the GPUs. Install Docker from Ubuntu's archive, add NVIDIA's repository, and configure Docker to use the NVIDIA runtime:

```
sudo apt install -y docker.io
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
sudo docker run --rm --gpus all ubuntu nvidia-smi
```

The last command should print the same table as on the host. Docker CE from Docker's own repository works the same way if you prefer it. As an example, this runs vLLM in a container and publishes it on localhost only:

```
sudo docker run --rm --gpus all --ipc=host -p 127.0.0.1:8000:8000 -v ~/.cache/huggingface:/root/.cache/huggingface vllm/vllm-openai:latest --model Qwen/Qwen2.5-7B-Instruct
```

> Docker writes its own firewall rules. A port published as `-p 8000:8000` is reachable from the internet even when ufw denies it. Publish ports as `127.0.0.1:PORT:PORT` unless they are meant to be public.

## Keep endpoints private

None of these tools is safe to expose as it is: Ollama and ComfyUI have no login at all, and vLLM's token is optional. Keep them bound to `127.0.0.1` and reach them through an SSH tunnel from your own computer (on Windows, the same command works in PowerShell):

```
ssh -N -L 8188:127.0.0.1:8188 -L 8000:127.0.0.1:8000 -L 11434:127.0.0.1:11434 YOUR_USER@SERVER_IP
```

While the tunnel runs, open `http://127.0.0.1:8188` in your browser for ComfyUI, and point OpenAI-compatible clients at `http://127.0.0.1:8000/v1`. Close the rest of the server with a firewall that allows only SSH, as in the [hardening guide](https://offshoreserv.com/docs/security/hardening). If a service must be public, put a reverse proxy with TLS, authentication and rate limits in front of it, and keep the [acceptable use policy](https://offshoreserv.com/acceptable-use-policy) in mind.

## Monitor the GPUs

```
nvidia-smi dmon -s pucm
watch -n 1 nvidia-smi
nvidia-smi --query-gpu=index,temperature.gpu,utilization.gpu,memory.used,memory.total,power.draw --format=csv -l 5
sudo apt install -y nvtop
nvtop
```

`nvidia-smi dmon` prints one line per GPU every second, with power and temperature (p), utilization (u), clocks (c) and memory (m). The query form writes CSV every five seconds, handy for logging long jobs. `nvtop` is an interactive view of all GPUs and their processes, similar to htop. If clocks drop under load, check why:

```
nvidia-smi -q -d PERFORMANCE
```

## Troubleshooting

- **"CUDA out of memory":** the model, its context or the batch does not fit. Choose a smaller or more compressed model, lower `--max-model-len` in vLLM, or reduce the batch size. `nvidia-smi` shows whether another process still holds memory.
- **`torch.cuda.is_available()` returns False:** the virtual environment is not active, or the PyTorch build expects a newer CUDA version than the driver supports. Compare `torch.version.cuda` with the CUDA version in `nvidia-smi`.
- **"Driver/library version mismatch":** the driver packages were updated but the old kernel module is still loaded. Reboot.
- **A container sees no GPU:** add `--gpus all` to `docker run`, and check that `nvidia-ctk runtime configure` ran and Docker was restarted.
- **Slow first start:** the first run downloads models and compiles GPU kernels. Later starts are faster.

Ready to pick hardware? Compare the [RTX servers](https://offshoreserv.com/offshore-gpu-servers/rtx), the [data center GPUs](https://offshoreserv.com/offshore-gpu-servers/datacenter) and the [multi-GPU servers](https://offshoreserv.com/offshore-gpu-servers/multi-gpu).

**Stuck on a step?**

Dedicated and GPU customers can open a ticket from the [client area](https://offshoreserv.com/account/support) with the server’s IP address and what they tried. First reply target: under 12 hours. For every other server, use the Server actions on its page, the [guides](https://offshoreserv.com/docs) and the [network status](https://offshoreserv.com/status) page.

---

OffshoreServ is an offshore hosting provider: VPS, dedicated servers, Windows RDP and GPU servers in seven jurisdictions (Iceland, Switzerland, Moldova, Romania, the Netherlands, Bulgaria and Malaysia), paid only in cryptocurrency (Bitcoin, Ethereum, Monero, Tether (USDT) and Solana), with no identity checks (no KYC).

Prices and plans: https://offshoreserv.com/pricing · Answers: https://offshoreserv.com/faq · Every page: https://offshoreserv.com/llms.txt
