
Multi-GPU servers.
Up to 320 GB of VRAM.
Two or four GPUs in one server for models that do not fit on a single card:
64 to 320 GB of VRAM- NVLink on H100 nodes
- Up to
30 TB of NVMe - Ready in
1 to 24 hours
01Plans & pricing
Three multi-GPU servers.
All GPUs in a server are dedicated to you, with large NVMe arrays for weights and datasets.
| GPU | VRAM | CPU | RAM | Storage | Locations | Price | Order |
|---|---|---|---|---|---|---|---|
|
|
$558.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$1,096.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$2,348.99/mo
Cheapest competitor |
Deploy |
No GPU plan in this jurisdiction yet: or compare them.
RTX 4090 , 5090, L40S, A100, H100- CUDA drivers preinstalled
- Single and
multi-GPU nodes No KYC · BTC, ETH, XMR, USDT, SOL
02Sizing
What fits in one node.
Dense models with an
| Server | Total VRAM | Largest model at | At | GPU link |
|---|---|---|---|---|
| 70B | 14B | PCIe | ||
| 123B | 32B (70B at | NVLink | ||
| 405B | 123B | NVLink |
Mixture-of-experts models such as
03How it works
One model, several GPUs.
Why the link between the GPUs matters as much as the GPUs themselves.
Serving frameworks split a model across GPUs in two ways. Tensor parallelism divides every layer between the cards and needs a fast link between them; pipeline or layer parallelism gives each card a block of layers and tolerates a slower link.
That is why the link matters. On our H100 servers the GPUs talk over NVLink, so tensor parallelism scales well for serving and training. The two
Run nvidia-smi topo -m--tensor-parallel-size
| Server | Memory | Bandwidth |
|---|---|---|
04Use cases
When one GPU is not enough.
Models and jobs that need the memory, or the throughput, of several GPUs in one box.
70B to 405B models
Serve the largest
open-weight models privately, split across GPUs.Distributed training
Data- and
tensor-parallel training over NVLink.Parallel jobs
One model
per GPU : several experiments or users at once.High-throughput servingMany concurrent users on one endpoint with batching.
05Jurisdictions
Where multi-GPU servers are available.
Moldova and the Netherlands host our
-
MoldovaChișinăuBest value Outside the EU
· no DSA London~33 ms New York ~113 ms Singapore~129 ms -
NetherlandsAmsterdamNetwork hub EU member
· DSA applies London~7 ms New York ~87 ms Singapore~154 ms - Compare all seven EU status, what can remove content and latency, side by side. Open the comparison
Latency figures are estimates from distance, not measurements. Every jurisdiction has the same prices. How we estimate latency
06Offshore, in writing
The rules, before you pay.
What we promise is written into our policies, not just our marketing.
US DMCA noticesNot actionedAnswered with our policy, never enforced. Only a local
DMCA policycourt order , or a valid EU notice in EU locations, can require action.IdentityEmail only
No name, address, phone or ID document, ever. A private or disposable address is fine.
Privacy policyPayment5 cryptocurrencies
Bitcoin, Ethereum, Monero, USDT and Solana, paid
Crypto paymentson-chain from any wallet to your balance. No card processor, no chargebacks.TransparencySigned canary
A
Warrant canaryPGP-signed warrant canary every quarter and a public count of every request we receive.
07FAQ
Multi-GPU servers, answered.
Another question? The full FAQ answers what people ask before an order: payments, privacy, complaints and support.
Read the full FAQCan I use the GPUs separately?
Yes. Each GPU is visible to the system on its own: run one model CUDA_VISIBLE_DEVICES
Is NVLink included on every multi-GPU server?
On the
Can I get eight GPUs or several nodes?
Yes, as a custom build. Once your first server is delivered, describe the cluster you need in a ticket from your client area and we reply with a price.
How fast is delivery?
Between 1 and
Can I get a refund?
GPU servers are not refundable once delivered, because the hardware is set aside for you.
The largest open models, on your own node.
- 1Pick 2 or
4 GPUs - 2Pay in crypto,
no ID asked - 3SSH in within
1 to 24 hours