Launch pricing: every plan costs 30% less than the cheapest offshore competitor we track. See the benchmarkEvery plan 30% under the cheapest offshore host

2 × RTX 5090 · 2 × H100 · 4 × H100

Multi-GPU servers.
Up to 320 GB of VRAM.

Two or four GPUs in one server for models that do not fit on a single card: 2 × RTX 5090 with 64 GB, or 2 × and 4 × H100 with NVLink and up to 320 GB of HBM3. From $558.99/mo, no KYC.

  • 64 to 320 GB of VRAM
  • NVLink on H100 nodes
  • Up to 30 TB of NVMe
  • Ready in 1 to 24 hours

Configure your multi-GPU serverReady in 1–24 h

JurisdictionMoldova · Chișinău

PlanCompare plans

Billing cycle

$1,096.99/mo

Monthly · cancel anytime

Cheapest competitor $1,567.50−30%

No IDBTC, XMR, USDT +2No setup fee

01Plans & pricing

Three multi-GPU servers.

All GPUs in a server are dedicated to you, with large NVMe arrays for weights and datasets.

GPU Servers: 3 plans, monthly prices in USD
GPU VRAM CPU RAM Storage Locations Price Order
2 × RTX 5090 64 GB of VRAM in one box 64 GB VRAMEPYC 9354 32c256 GB RAM2 × 1.92 TB NVMeMoldovaNetherlands 64 GB EPYC 9354 · 32c 256 GB 2 × 1.92 TB NVMe MoldovaNetherlands $558.99/mo Cheapest competitor $799.00 Deploy 2 × RTX 5090
2 × H100 160 GB HBM3, NVLink 160 GB VRAM48 vCPU384 GB RAM4 TB NVMeMoldovaNetherlands 160 GB 48 vCPU 384 GB 4 TB NVMe MoldovaNetherlands $1,096.99/mo Cheapest competitor $1,567.50 Deploy 2 × H100
  • RTX 4090, 5090, L40S, A100, H100
  • CUDA drivers preinstalled
  • Single and multi-GPU nodes
  • No KYC · BTC, ETH, XMR, USDT, SOL

02Sizing

What fits in one node.

Dense models with an 8K-token context and 10% headroom, split across the GPUs by vLLM or llama.cpp. Longer contexts need more.

What fits in one node.
ServerTotal VRAMLargest model at 4-bitAt 16-bitGPU link
2 × RTX 509064 GB70B14BPCIe
2 × H100160 GB123B32B (70B at 8-bit)NVLink
4 × H100320 GB405B123BNVLink

Mixture-of-experts models such as gpt-oss-120b need less compute per token but still need all their weights in memory.

03How it works

One model, several GPUs.

Why the link between the GPUs matters as much as the GPUs themselves.

Serving frameworks split a model across GPUs in two ways. Tensor parallelism divides every layer between the cards and needs a fast link between them; pipeline or layer parallelism gives each card a block of layers and tolerates a slower link.

That is why the link matters. On our H100 servers the GPUs talk over NVLink, so tensor parallelism scales well for serving and training. The two RTX 5090 cards talk over PCIe: excellent for inference of a model split in two, more limited for training.

Run nvidia-smi topo -m after delivery to see the topology, then set --tensor-parallel-size in vLLM to the number of GPUs. The GPU images guide has the full commands.

Multi-GPU servers: memory and bandwidth
ServerMemoryBandwidth per GPU
2 × RTX 50902 × 32 GB GDDR71,792 GB/s
2 × H1002 × 80 GB HBM33.35 TB/s
4 × H1004 × 80 GB HBM33.35 TB/s

04Use cases

When one GPU is not enough.

Models and jobs that need the memory, or the throughput, of several GPUs in one box.

  • 70B to 405B models

    Serve the largest open-weight models privately, split across GPUs.

    4 × H100
  • Distributed training

    Data- and tensor-parallel training over NVLink.

    2 × H100
  • Parallel jobs

    One model per GPU: several experiments or users at once.

    2 × RTX 5090
  • High-throughput serving

    Many concurrent users on one endpoint with batching.

    vLLM guide

06Offshore, in writing

The rules, before you pay.

What we promise is written into our policies, not just our marketing.

  • US DMCA noticesNot actioned

    Answered with our policy, never enforced. Only a local court order, or a valid EU notice in EU locations, can require action.

    DMCA policy
  • IdentityEmail only

    No name, address, phone or ID document, ever. A private or disposable address is fine.

    Privacy policy
  • Payment5 cryptocurrencies

    Bitcoin, Ethereum, Monero, USDT and Solana, paid on-chain from any wallet to your balance. No card processor, no chargebacks.

    Crypto payments
  • TransparencySigned canary

    A PGP-signed warrant canary every quarter and a public count of every request we receive.

    Warrant canary

07FAQ

Multi-GPU servers, answered.

Another question? The full FAQ answers what people ask before an order: payments, privacy, complaints and support.

Read the full FAQ
Can I use the GPUs separately?

Yes. Each GPU is visible to the system on its own: run one model per card with CUDA_VISIBLE_DEVICES, or split one model across all of them.

Is NVLink included on every multi-GPU server?

On the 2 × and 4 × H100 servers, yes. The RTX 5090 does not support NVLink, so the 2 × RTX 5090 server connects its cards over PCIe.

Can I get eight GPUs or several nodes?

Yes, as a custom build. Once your first server is delivered, describe the cluster you need in a ticket from your client area and we reply with a price.

How fast is delivery?

Between 1 and 24 hours after you order. Multi-GPU nodes depend on stock: if one is unavailable in your jurisdiction, we tell you and either hold your order or return the payment to your balance.

Can I get a refund?

GPU servers are not refundable once delivered, because the hardware is set aside for you.

The largest open models, on your own node.

  1. 1Pick 2 or 4 GPUs
  2. 2Pay in crypto, no ID asked
  3. 3SSH in within 1 to 24 hours
Configure your server All GPU servers

From $558.99/mo · dedicated GPUs · NVLink on H100

Welcome back

Sign in to manage your servers and your balance.

No KYCHuman check by Cloudflare TurnstileNo tracking