RTX 4090 Best value
- Memory bandwidth
1,008 GB/s - Fits at
16-bit - Up to about 8B
- Fits at
8-bit - Up to about 14B
- Image models
- SDXL, Flux
- Best for
- Chat assistants, image generation

RTX cards give the most inference speed per dollar:
01Plans & pricing
Each card is dedicated to you. Every server has NVMe storage and an unmetered
| GPU | VRAM | CPU | RAM | Storage | Locations | Price | Order |
|---|---|---|---|---|---|---|---|
|
|
$61.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$84.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$134.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$306.99/mo
Cheapest competitor |
Deploy |
|||||
|
|
$558.99/mo
Cheapest competitor |
Deploy |
No GPU plan in this jurisdiction yet: or compare them.
Need HBM memory or NVLink? See A100, L40S & H100.
02Choose
The 4090 is the value pick. The 5090 adds
03Sizing
Memory a dense model needs with an
| Model size | |||
|---|---|---|---|
| 7–8B | ~ | ~ | ~ |
| 13–14B | ~ | ~ | ~ |
| 32B | ~ | ~ | ~ A100 or H100 |
| 70B | ~ A100, H100 or | ~ | ~ |
A
04Use cases
Consumer and workstation cards shine where a model fits in
8B to 32B models behind Open WebUI, for you or your team.
Stable Diffusion and Flux workflows in ComfyUI.
Video models and Blender Cycles renders on a card you do not share.
05Jurisdictions
Each card has its own list; the plans above show exactly where.
Latency figures are estimates from distance, not measurements. Every jurisdiction has the same prices. How we estimate latency
06Offshore, in writing
What we promise is written into our policies, not just our marketing.
Answered with our policy, never enforced. Only a local
IdentityEmail only
No name, address, phone or ID document, ever. A private or disposable address is fine.
Privacy policyPayment5 cryptocurrencies
Bitcoin, Ethereum, Monero, USDT and Solana, paid
TransparencySigned canary
A
07FAQ
Another question? The full FAQ answers what people ask before an order: payments, privacy, complaints and support.
Read the full FAQFor inference of models that fit in
Yes. vLLM and llama.cpp split a model across both cards (tensor or layer parallelism), which gives you
If you need
Yes, on the
GPU servers are not refundable once delivered, because the hardware is set aside for you. Check the specifications and the GPU images guide before you order.