Solutions · Inference serving
GPU servers for inference hosting: always on, flat price
An endpoint that has to answer at 3 a.m. is the worst possible fit for hourly billing: the meter runs at full rate whether anyone is calling. A dedicated server at a flat monthly price costs the same at 3 a.m. as at peak — and the tokens are free.
- Recommended GPU$429NVIDIA L40S, 48 GB GDDR6, per month
- Same GPU, hourly on-demand$1,124market median × 730 h
- You keep−62%$695 a month, every month
Hardware
Which GPU for inference serving
Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.
| GPU | Why for this workload | Per GPU, per month | Availability |
|---|---|---|---|
| NVIDIA L424 GB GDDR6 · Ada Lovelace | The smallest always-on endpoint: 7B–8B models at FP8, Whisper, embeddings and rerankers. | $169 | 26 left |
| NVIDIA L40S48 GB GDDR6 · Ada Lovelace | Our recommendation: 48 GB with FP8 for 30B-class models, plus video engines for multimodal work. | $429 | 17 left |
| NVIDIA RTX PRO 6000 Blackwell96 GB GDDR7 ECC · Blackwell | 96 GB at a workstation price: 70B-class models at FP8 on one GPU. | $749 | 6 left |
| NVIDIA H200141 GB HBM3e · Hopper | Bandwidth and 141 GB for high-throughput serving of 70B models with long contexts. | $1,690 | 9 left |
Highlighted row: our default recommendation for inference serving. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.
Sizing
How much memory does a model need to serve?
Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.
From order to first run
- Size the model to the memoryFP8 or 4-bit weights, then KV cache for your longest prompt times your concurrency — the table above is the whole calculation.
- Boot a serving imagevLLM 0.9 with an OpenAI-compatible API, NVIDIA Triton with TensorRT, or Ollama with Open WebUI — chosen at checkout.
- Put it on your IPA dedicated public IPv4 and a /64 IPv6 block on every server, reverse DNS on request, DDoS mitigation always on.
- Forget the meterThe renewal settles from your prepaid balance every month; traffic spikes, long prompts and big batches do not change the invoice.
In every server
- Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
- 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
- Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
- Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
- Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
- DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Inference serving on a monthly GPU: questions answered
How much does it cost to host an LLM endpoint?
A 7B–8B model runs on an L4 at $169 a month; a 30B-class model at FP8 on an L40S at $429; a 70B model at FP8 on an RTX PRO 6000 at $749 or an H200 at $1,690. Every price includes the server, 20 TB of transfer and 24/7 support — there is no per-token or per-request charge.
Do you offer serverless or per-token pricing?
No. We rent the whole machine by the month. For traffic that is bursty and idle most of the day, a serverless provider will be cheaper; for a steady endpoint, the flat rate beats any meter — and the latency is yours to control.
Can I run several models on one server?
Yes — it is your machine. Run several vLLM instances, use MIG on an A100 to isolate them, or run Triton with multiple backends. Memory is the only limit.
Start inference serving tonight
Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.