
NVIDIA
Offshore GPU servers.
Private AI, no KYC .
Dedicated NVIDIA GPUs from
- Ready in
1 to 24 hours - Dedicated GPUs, never shared
- Prompts stay on your server
No KYC , email only
01Plans & pricing
10 GPU servers, from RTX A4000 to 4 × H100.
Every GPU is dedicated to you: no
| GPU | VRAM | CPU | RAM | Storage | Locations | Price | Order |
|---|---|---|---|---|---|---|---|
|
|
$61.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$84.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$134.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$306.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$355.99/mo
Cheapest competitor |
Deploy |
|||||
|
L40S
Inference at scale
48 GB VRAM |
$418.99/mo
Cheapest competitor |
Deploy L40S | |||||
|
|
$558.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$581.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$1,096.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$2,348.99/mo
Cheapest competitor |
Deploy |
No GPU plan in this jurisdiction yet: or compare them.
RTX 4090 , 5090, L40S, A100, H100- CUDA drivers preinstalled
- Single and
multi-GPU nodes No KYC · BTC, ETH, XMR, USDT, SOL
Planning a cluster or a model that needs more than
02Families
Three kinds of GPU server.
Consumer RTX cards for price per token, datacenter GPUs for memory and bandwidth,
RTX A4000 to 6000 Ada16 to 64 GB of VRAM. The most inference per dollar, for 7B to 70B models and image generation.From $61.99/moSee the servers- A100, L40S & H100HBM or ECC memory and up to
3.35 TB/s of bandwidth, for training and serving at scale.From $355.99/moSee the servers Multi-GPU nodesTwo or four GPUs in one server, up to320 GB of VRAM with NVLink on H100.From $558.99/moSee the servers- Custom buildAlready a GPU or dedicated customer? Ask by ticket for another GPU, more memory or storage, or a specific jurisdiction, and we reply with a price.How support works
03By model
Rent a specific GPU.
The four cards people ask for most, each with its specifications, what fits in its memory and its price per hour of use.
RTX 4090 ,24 GB The value card for private chat assistants and image generation.$84.99/moRTX 4090 serversRTX 5090 ,32 GB Blackwell and GDDR7: faster tokens, longer contexts, or two cards for 70B.$134.99/moRTX 5090 servers- A100,
80 GB 80 GB of HBM2e on one card, forfine-tuning and large models.$355.99/moA100 servers - H100,
80 GB HBM3 at3.35 TB/s and FP8, alone or up to four with NVLink.$581.99/moH100 servers
04Private LLM hosting
Your models, your prompts, your server.
Run
Llama, Qwen, Mistral, DeepSeek, Gemma and other
The table gives the memory a model needs at each precision with an
- Chat with a private assistant from your browser through an SSH tunnel.
- Expose an API to your own apps, never to a
third-party provider. Fine-tune with LoRA on your own data, which never leaves the server.
Set up Ollama, vLLM or ComfyUI · How much VRAM do LLMs need?
| Model size | |||
|---|---|---|---|
| 7–8B | ~ | ~ | ~ |
| 13–14B | ~ | ~ | ~ |
| 32B | ~ | ~ | ~ A100 or H100 |
| 70B | ~ A100, H100 or | ~ | ~ |
| 123B | ~ | ~ | ~ |
05Included
Ready for CUDA on first login.
Check the card, pull a model, start serving.
Dedicated GPUs
Each card is yours alone: no
time-slicing , no shared memory.Driver & CUDA ready
Ubuntu 24.04 with the NVIDIA driver and the CUDA toolkit installed.A guide for every stackStep by step for PyTorch, vLLM, Ollama, ComfyUI and Docker.
NVMe storage
1 to 30 TB of NVMe for datasets, checkpoints and weights.1 Gbps , unmeteredPull models and datasets without a traffic meter.
Nothing logged
We never see or log your prompts, outputs or training data.
DDoS mitigation
Public endpoints stay online under attack.
Full root
Your drivers, your kernel modules, your stack.
06Compare
Every GPU, side by side.
For inference, memory bandwidth sets the tokens per second; VRAM sets which models fit.
| Specification | L40S | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| GPU | ||||||||||
| VRAM | ||||||||||
| Memory type | GDDR6 ECC | GDDR6X | GDDR7 | GDDR6 ECC | HBM2e | GDDR6 ECC | GDDR7 | HBM3 | HBM3 | HBM3 |
| Bandwidth | ~ | |||||||||
| CPU | ||||||||||
| System memory | ||||||||||
| Storage | ||||||||||
| Jurisdictions | ||||||||||
| Monthly | $61.99 | $84.99 | $134.99 | $306.99 | $355.99 | $418.99 | $558.99 | $581.99 | $1,096.99 | $2,348.99 |
| Quarterly, | $58.89 | $80.74 | $128.24 | $291.64 | $338.19 | $398.04 | $531.04 | $552.89 | $1,042.14 | $2,231.54 |
| Yearly, | $54.55 | $74.79 | $118.79 | $270.15 | $313.27 | $368.71 | $491.91 | $512.15 | $965.35 | $2,067.11 |
| Cheapest competitor | ||||||||||
| Order | Deploy | Deploy | Deploy | Deploy | Deploy | Deploy L40S | Deploy | Deploy | Deploy | Deploy |
07Use cases
What people run on offshore GPUs.
From a private chat assistant to
LLM inference
Private chat and APIs on
open-weight models, from 8B to 405B.Fine-tuning LoRA and QLoRA on your own data, which never leaves the server.
Image & video generation
Stable Diffusion, Flux and video models with ComfyUI.
3D rendering
Blender Cycles and other GPU renderers, without tying up your workstation.
Speech & transcription
Whisper and
text-to-speech models on audio that must stay private.Research & data science
Jupyter, PyTorch and CUDA experiments on dedicated hardware.
08Jurisdictions
Four jurisdictions for GPUs.
Availability depends on the card: the plan table shows where each one is in stock.
-
MoldovaChișinăuBest value Outside the EU
· no DSA London~33 ms New York ~113 ms Singapore~129 ms -
NetherlandsAmsterdamNetwork hub EU member
· DSA applies London~7 ms New York ~87 ms Singapore~154 ms -
IcelandReykjavík
Free-speech haven EEA· outside the EU· no DSA London~29 ms New York ~63 ms Singapore~169 ms -
RomaniaBucharestBudget EU EU member
· DSA applies London~32 ms New York ~113 ms Singapore~131 ms
Latency figures are estimates from distance, not measurements. Every jurisdiction has the same prices. How we estimate latency
09Offshore, in writing
The rules, before you pay.
What we promise is written into our policies, not just our marketing.
US DMCA noticesNot actionedAnswered with our policy, never enforced. Only a local
DMCA policycourt order , or a valid EU notice in EU locations, can require action.IdentityEmail only
No name, address, phone or ID document, ever. A private or disposable address is fine.
Privacy policyPayment5 cryptocurrencies
Bitcoin, Ethereum, Monero, USDT and Solana, paid
Crypto paymentson-chain from any wallet to your balance. No card processor, no chargebacks.TransparencySigned canary
A
Warrant canaryPGP-signed warrant canary every quarter and a public count of every request we receive.
10FAQ
GPU servers, answered.
Another question? The full FAQ answers what people ask before an order: payments, privacy, complaints and support.
Read the full FAQAre the GPUs shared or time-sliced ?
No. Every GPU in your plan is dedicated to your server for as long as you rent it. Nobody else runs jobs on it, and its memory is never shared.
Which GPU do I need for my model?
Start from the memory the weights need: see the table above or our guide on VRAM for LLMs. As a rule, a
Are the NVIDIA driver and CUDA installed?
Yes. The Ubuntu images ship with the NVIDIA driver and the CUDA toolkit. PyTorch, vLLM, Ollama and ComfyUI install in a few commands, as shown in the GPU images guide.
Can I run a public AI service?
Yes, as long as it is legal in your server’s jurisdiction and follows our acceptable use policy. You are responsible for what your service generates and serves.
How fast is delivery?
Between 1 and
Can I mine cryptocurrency on a GPU server?
Yes. Mining is allowed on GPU and dedicated servers, which are provisioned for you alone.
Can I get a refund?
GPU servers are not refundable once delivered, because the hardware is set aside for you. Check the specifications and the GPU images guide before ordering; your balance can always pay for other services.
Do you log prompts, outputs or datasets?
No. We do not log or inspect your server’s traffic, and we never look at what runs inside it.
Your own GPU, away from the big clouds.
- 1Pick a GPU and a jurisdiction
- 2Pay in crypto,
no ID asked - 3SSH in within
1 to 24 hours
From $61.99/mo