Launch pricing: every plan costs 30% less than the cheapest offshore competitor we track. See the benchmarkEvery plan 30% under the cheapest offshore host

RTX A4000 · 4090 · 5090 · 6000 Ada

RTX 4090 & 5090 servers.
The best price per token.

RTX cards give the most inference speed per dollar: 24 GB on the 4090, 32 GB of GDDR7 on the 5090 and 48 GB on the RTX 6000 Ada. Dedicated cards, driver and CUDA preinstalled, paid in crypto. From $61.99/mo.

  • 16 to 64 GB of VRAM
  • Up to 1.8 TB/s per card
  • Private inference
  • Ready in 1 to 24 hours

Configure your RTX serverReady in 1–24 h

JurisdictionMoldova · Chișinău

PlanCompare plans

Billing cycle

$84.99/mo

Monthly · cancel anytime

Cheapest competitor $121.50−30%

No IDBTC, XMR, USDT +2No setup fee

01Plans & pricing

Five RTX servers, 16 to 64 GB of VRAM.

Each card is dedicated to you. Every server has NVMe storage and an unmetered 1 Gbps port.

GPU Servers: 5 plans, monthly prices in USD
GPU VRAM CPU RAM Storage Locations Price Order
RTX A4000 Entry inference, 7-13B models 16 GB VRAM8 vCPU64 GB RAM1 TB NVMeMoldovaNetherlandsRomania 16 GB 8 vCPU 64 GB 1 TB NVMe MoldovaNetherlandsRomania $61.99/mo Cheapest competitor $89.00 Deploy RTX A4000
RTX 5090 Fast inference, image and video 32 GB VRAM12 vCPU96 GB RAM2 TB NVMeMoldovaNetherlandsIceland 32 GB 12 vCPU 96 GB 2 TB NVMe MoldovaNetherlandsIceland $134.99/mo Cheapest competitor $193.00 Deploy RTX 5090
RTX 6000 Ada 48 GB, 32B at 8-bit 48 GB VRAMEPYC 7302P 16c128 GB RAM2 × 1.92 TB NVMeMoldovaNetherlands 48 GB EPYC 7302P · 16c 128 GB 2 × 1.92 TB NVMe MoldovaNetherlands $306.99/mo Cheapest competitor $439.00 Deploy RTX 6000 Ada
2 × RTX 5090 64 GB of VRAM in one box 64 GB VRAMEPYC 9354 32c256 GB RAM2 × 1.92 TB NVMeMoldovaNetherlands 64 GB EPYC 9354 · 32c 256 GB 2 × 1.92 TB NVMe MoldovaNetherlands $558.99/mo Cheapest competitor $799.00 Deploy 2 × RTX 5090
  • RTX 4090, 5090, L40S, A100, H100
  • CUDA drivers preinstalled
  • Single and multi-GPU nodes
  • No KYC · BTC, ETH, XMR, USDT, SOL

Need HBM memory or NVLink? See A100, L40S & H100.

02Choose

RTX 4090 or RTX 5090?

The 4090 is the value pick. The 5090 adds 8 GB and about 78% more memory bandwidth, which shows in tokens per second.

RTX 4090 Best value

24 GB GDDR6X

Memory bandwidth
1,008 GB/s
Fits at 16-bit
Up to about 8B
Fits at 8-bit
Up to about 14B
Image models
SDXL, Flux
Best for
Chat assistants, image generation

RTX 5090

32 GB GDDR7

Memory bandwidth
1,792 GB/s
Fits at 4-bit
Up to about 32B
Fits at 8-bit
14B with long context
Image models
Flux, video models
Best for
Faster tokens, longer context

03Sizing

What fits on each card.

Memory a dense model needs with an 8K-token context and 10% headroom. Longer contexts need more.

What fits on each card.
Model size4-bit8-bit16-bit
7–8B~7 GB
RTX A4000
~11 GB
RTX A4000
~19 GB
RTX 4090
13–14B~11 GB
RTX A4000
~18 GB
RTX 4090
~32 GB
RTX 6000 Ada, L40S
32B~24 GB
RTX 5090
~40 GB
RTX 6000 Ada, L40S
~73 GB
A100 or H100
70B~50 GB
A100, H100 or 2 × RTX 5090
~85 GB
2 × H100
~157 GB
4 × H100

A 4-bit 70B model needs about 50 GB: 2 × RTX 5090 (64 GB), or an 80 GB A100 or H100. Full guide: how much VRAM do LLMs need?

04Use cases

What RTX servers are best at.

Consumer and workstation cards shine where a model fits in 24 to 64 GB of VRAM.

  • Private chat assistants

    8B to 32B models behind Open WebUI, for you or your team.

    Ollama guide
  • Image generation

    Stable Diffusion and Flux workflows in ComfyUI.

    RTX 5090
  • Video & 3D

    Video models and Blender Cycles renders on a card you do not share.

    2 × RTX 5090
  • Coding assistants

    Self-hosted code models answering from your own repositories.

    vLLM guide

06Offshore, in writing

The rules, before you pay.

What we promise is written into our policies, not just our marketing.

  • US DMCA noticesNot actioned

    Answered with our policy, never enforced. Only a local court order, or a valid EU notice in EU locations, can require action.

    DMCA policy
  • IdentityEmail only

    No name, address, phone or ID document, ever. A private or disposable address is fine.

    Privacy policy
  • Payment5 cryptocurrencies

    Bitcoin, Ethereum, Monero, USDT and Solana, paid on-chain from any wallet to your balance. No card processor, no chargebacks.

    Crypto payments
  • TransparencySigned canary

    A PGP-signed warrant canary every quarter and a public count of every request we receive.

    Warrant canary

07FAQ

RTX servers, answered.

Another question? The full FAQ answers what people ask before an order: payments, privacy, complaints and support.

Read the full FAQ
Why choose RTX over an A100 or H100?

For inference of models that fit in 24 to 48 GB, RTX cards deliver more tokens per dollar. Datacenter GPUs make sense when you need 80 GB on one card, HBM bandwidth for training, or NVLink between GPUs.

Can two RTX 5090s run one model?

Yes. vLLM and llama.cpp split a model across both cards (tensor or layer parallelism), which gives you 64 GB for a 4-bit 70B model. The cards talk over PCIe, not NVLink, so scaling is good for inference and more limited for training.

Is the RTX 6000 Ada worth it?

If you need 48 GB on a single card with ECC memory, yes: it runs 32B models at 8-bit and 14B models at full precision without splitting them across GPUs. A 4-bit 70B model needs about 50 GB, so plan for two cards or 80 GB. For smaller models, the RTX 5090 is faster and cheaper.

Are the driver and CUDA installed?

Yes, on the Ubuntu 24.04 image. The RTX 5090 needs a recent CUDA build of PyTorch; the GPU images guide gives the exact commands.

Can I get a refund?

GPU servers are not refundable once delivered, because the hardware is set aside for you. Check the specifications and the GPU images guide before you order.

The fastest tokens per dollar, offshore.

  1. 1Pick a card and a jurisdiction
  2. 2Pay in crypto, no ID asked
  3. 3SSH in within 1 to 24 hours
Configure your RTX server All GPU servers

From $61.99/mo · dedicated cards · CUDA preinstalled

Welcome back

Sign in to manage your servers and your balance.

No KYCHuman check by Cloudflare TurnstileNo tracking