Launch

Launch pricing: up to 74% below on-demand, fixed for your term — ends Oct 31, 2026 · 42d left Launch pricing ends Oct 31 · 42d left

See pricing

Solutions · Inference serving

GPU servers for inference hosting: always on, flat price

An endpoint that has to answer at 3 a.m. is the worst possible fit for hourly billing: the meter runs at full rate whether anyone is calling. A dedicated server at a flat monthly price costs the same at 3 a.m. as at peak — and the tokens are free.

  • Recommended GPU$429NVIDIA L40S, 48 GB GDDR6, per month
  • Same GPU, hourly on-demand$1,124market median × 730 h
  • You keep−62%$695 a month, every month
  • Single-tenant bare metal, root access
  • Live in under 10 minutes after payment
  • No KYC · BTC, ETH, USDT
  • 99.9% uptime SLA · 24/7 engineers

Hardware

Which GPU for inference serving

Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.

GPUWhy for this workloadPer GPU, per monthAvailability
NVIDIA L424 GB GDDR6 · Ada Lovelace The smallest always-on endpoint: 7B–8B models at FP8, Whisper, embeddings and rerankers. $169 26 left
NVIDIA L40S48 GB GDDR6 · Ada Lovelace Our recommendation: 48 GB with FP8 for 30B-class models, plus video engines for multimodal work. $429 17 left
NVIDIA RTX PRO 6000 Blackwell96 GB GDDR7 ECC · Blackwell 96 GB at a workstation price: 70B-class models at FP8 on one GPU. $749 6 left
NVIDIA H200141 GB HBM3e · Hopper Bandwidth and 141 GB for high-throughput serving of 70B models with long contexts. $1,690 9 left

Highlighted row: our default recommendation for inference serving. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.

Sizing

How much memory does a model need to serve?

Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.

Weights, 16-bit2 bytes per parameter8B = 16 GB, 30B = 60 GB, 70B = 140 GB.
Weights, FP81 byte per parameter30B = 30 GB (L40S), 70B = 70 GB (RTX PRO 6000, H200).
Weights, 4-bit~0.5 bytes per parameter13B on an L4, 30B on a 4090, 70B on an L40S.
KV cachecontext × concurrencyLlama-70B: about 0.3 MB per token at 16-bit — a 128k-token context is ~42 GB. Budget for it.

From order to first run

  1. Size the model to the memoryFP8 or 4-bit weights, then KV cache for your longest prompt times your concurrency — the table above is the whole calculation.
  2. Boot a serving imagevLLM 0.9 with an OpenAI-compatible API, NVIDIA Triton with TensorRT, or Ollama with Open WebUI — chosen at checkout.
  3. Put it on your IPA dedicated public IPv4 and a /64 IPv6 block on every server, reverse DNS on request, DDoS mitigation always on.
  4. Forget the meterThe renewal settles from your prepaid balance every month; traffic spikes, long prompts and big batches do not change the invoice.

In every server

  • Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
  • 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
  • Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
  • Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
  • Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
  • DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Read the documentation

Inference serving on a monthly GPU: questions answered

How much does it cost to host an LLM endpoint?

A 7B–8B model runs on an L4 at $169 a month; a 30B-class model at FP8 on an L40S at $429; a 70B model at FP8 on an RTX PRO 6000 at $749 or an H200 at $1,690. Every price includes the server, 20 TB of transfer and 24/7 support — there is no per-token or per-request charge.

Do you offer serverless or per-token pricing?

No. We rent the whole machine by the month. For traffic that is bursty and idle most of the day, a serverless provider will be cheaper; for a steady endpoint, the flat rate beats any meter — and the latency is yours to control.

Can I run several models on one server?

Yes — it is your machine. Run several vLLM instances, use MIG on an A100 to isolate them, or run Triton with multiple backends. Memory is the only limit.

Start inference serving tonight

Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.