---
title: "Multi-GPU Servers: 2× RTX 5090, 2× & 4× H100 · No KYC"
description: "Two or four dedicated GPUs in one offshore server: 2 × RTX 5090, or 2 × and 4 × H100 with NVLink and up to 320 GB of VRAM."
url: https://offshoreserv.com/offshore-gpu-servers/multi-gpu
lang: en
updated: 2026-09-27
source: HTML page at the url above (canonical); this is its Markdown version
---

2 × RTX 5090 · 2 × H100 · 4 × H100

# Multi-GPU servers. Up to 320 GB of VRAM.

Two or four GPUs in one server for models that do not fit on a single card: 2 × RTX 5090 with 64 GB, or 2 × and 4 × H100 with NVLink and up to 320 GB of HBM3. **From $558.99/mo, no KYC.**

- 64 to 320 GB of VRAM
- NVLink on H100 nodes
- Up to 30 TB of NVMe
- Ready in 1 to 24 hours

## Three multi-GPU servers.

All GPUs in a server are dedicated to you, with large NVMe arrays for weights and datasets.

| GPU | VRAM | CPU | RAM | Storage | Price |
| --- | --- | --- | --- | --- | --- |
| **2 × RTX 5090** (64 GB of VRAM in one box) | 64 GB | EPYC 9354 · 32c | 256 GB | 2 × 1.92 TB NVMe | **$558.99**/mo |
| **2 × H100** (160 GB HBM3, NVLink) | 160 GB | 48 vCPU | 384 GB | 4 TB NVMe | **$1,096.99**/mo |
| **4 × H100** (320 GB HBM3, NVLink), cluster | 320 GB | 96 vCPU | 768 GB | 30 TB NVMe | **$2,348.99**/mo |

## What fits in one node.

Dense models with an 8K-token context and 10% headroom, split across the GPUs by vLLM or llama.cpp. Longer contexts need more.

| Server | Total VRAM | Largest model at 4-bit | At 16-bit | GPU link |
| --- | --- | --- | --- | --- |
| 2 × RTX 5090 | 64 GB | 70B | 14B | PCIe |
| 2 × H100 | 160 GB | 123B | 32B (70B at 8-bit) | NVLink |
| 4 × H100 | 320 GB | 405B | 123B | NVLink |

Mixture-of-experts models such as gpt-oss-120b need less compute per token but still need all their weights in memory.

## One model, several GPUs.

Why the link between the GPUs matters as much as the GPUs themselves.

Serving frameworks split a model across GPUs in two ways. **Tensor parallelism** divides every layer between the cards and needs a fast link between them; **pipeline or layer parallelism** gives each card a block of layers and tolerates a slower link.

That is why the link matters. On our H100 servers the GPUs talk over **NVLink**, so tensor parallelism scales well for serving and training. The two RTX 5090 cards talk over PCIe: excellent for inference of a model split in two, more limited for training.

Run `nvidia-smi topo -m` after delivery to see the topology, then set `--tensor-parallel-size` in vLLM to the number of GPUs. The [GPU images guide](https://offshoreserv.com/docs/gpu/images) has the full commands.

| Server | Memory | Bandwidth per GPU |
| --- | --- | --- |
| 2 × RTX 5090 | 2 × 32 GB GDDR7 | 1,792 GB/s |
| 2 × H100 | 2 × 80 GB HBM3 | 3.35 TB/s |
| 4 × H100 | 4 × 80 GB HBM3 | 3.35 TB/s |

## When one GPU is not enough.

Models and jobs that need the memory, or the throughput, of several GPUs in one box.

- **70B to 405B models** — Serve the largest open-weight models privately, split across GPUs. — 4 × H100
- **Distributed training** — Data- and tensor-parallel training over NVLink. — 2 × H100
- **Parallel jobs** — One model per GPU: several experiments or users at once. — 2 × RTX 5090
- **High-throughput serving** — Many concurrent users on one endpoint with batching. — [vLLM guide](https://offshoreserv.com/docs/gpu/images)

## Where multi-GPU servers are available.

Moldova and the Netherlands host our multi-GPU nodes.

- [Moldova Chișinău Best value — Outside the EU · no DSA — London **~33 ms** New York **~113 ms** Singapore **~129 ms**](https://offshoreserv.com/locations/moldova)
- [Netherlands Amsterdam Network hub — EU member · DSA applies — London **~7 ms** New York **~87 ms** Singapore **~154 ms**](https://offshoreserv.com/locations/netherlands)
- [Compare all seven — EU status, what can remove content and latency, side by side.](https://offshoreserv.com/locations#compare)

Latency figures are estimates from distance, not measurements. Every jurisdiction has the same prices. [How we estimate latency](https://offshoreserv.com/network)

## The rules, before you pay.

What we promise is written into our policies, not just our marketing.

- US DMCA notices **Not actioned** — Answered with our policy, never enforced. Only a local court order, or a valid EU notice in EU locations, can require action.
- Identity **Email only** — No name, address, phone or ID document, ever. A private or disposable address is fine.
- Payment **5 cryptocurrencies** — Bitcoin, Ethereum, Monero, USDT and Solana, paid on-chain from any wallet to your balance. No card processor, no chargebacks.
- Transparency **Signed canary** — A PGP-signed warrant canary every quarter and a public count of every request we receive.

## Multi-GPU servers, answered.

**Another question?** The full FAQ answers what people ask before an order: payments, privacy, complaints and support.

### Can I use the GPUs separately?

Yes. Each GPU is visible to the system on its own: run one model per card with `CUDA_VISIBLE_DEVICES`, or split one model across all of them.

### Is NVLink included on every multi-GPU server?

On the 2 × and 4 × H100 servers, yes. The RTX 5090 does not support NVLink, so the 2 × RTX 5090 server connects its cards over PCIe.

### Can I get eight GPUs or several nodes?

Yes, as a custom build. Once your first server is delivered, describe the cluster you need in a ticket from your [client area](https://offshoreserv.com/account/support) and we reply with a price.

### How fast is delivery?

Between 1 and 24 hours after you order. Multi-GPU nodes depend on stock: if one is unavailable in your jurisdiction, we tell you and either hold your order or return the payment to your balance.

### Can I get a refund?

GPU servers are not refundable once delivered, because the hardware is set aside for you.

## Keep reading.

Guides, policies and articles picked for this page, written by our team.

- [H100 servers 80 GB of HBM3, up to 4 × H100.](https://offshoreserv.com/offshore-gpu-servers/h100)
- [RTX 5090 servers 32 GB of GDDR7 for $134.99/mo.](https://offshoreserv.com/offshore-gpu-servers/rtx-5090)
- [How much VRAM do LLMs need?Model sizes, quantization and the right GPU.](https://offshoreserv.com/blog/how-much-vram-for-llms)
- [AI-ready GPU images Drivers, CUDA and the tools preinstalled on GPU servers.](https://offshoreserv.com/docs/gpu/images)
- [A100, L40S & H100 HBM, ECC and NVLink for training.](https://offshoreserv.com/offshore-gpu-servers/datacenter)
- [RTX GPU servers The best price per token for inference.](https://offshoreserv.com/offshore-gpu-servers/rtx)

## The largest open models, on your own node.

1. Pick 2 or 4 GPUs
2. Pay in crypto, no ID asked
3. SSH in within 1 to 24 hours

From $558.99/mo · dedicated GPUs · NVLink on H100

---

OffshoreServ is an offshore hosting provider: VPS, dedicated servers, Windows RDP and GPU servers in seven jurisdictions (Iceland, Switzerland, Moldova, Romania, the Netherlands, Bulgaria and Malaysia), paid only in cryptocurrency (Bitcoin, Ethereum, Monero, Tether (USDT) and Solana), with no identity checks (no KYC).

Prices and plans: https://offshoreserv.com/pricing · Answers: https://offshoreserv.com/faq · Every page: https://offshoreserv.com/llms.txt
