Launch

Launch pricing: up to 74% below on-demand, fixed for your term — ends Oct 31, 2026 · 42d left Launch pricing ends Oct 31 · 42d left

See pricing

Solutions · Model training

GPU servers for model training, billed by the month

A training run is the definition of a machine that stays busy: weeks of wall-clock time at full utilisation, checkpoints every few hours, and a bill that hourly clouds price for the idle capacity they keep on standby. A monthly server or node removes the meter — and the idle.

  • Recommended GPU$1,249NVIDIA H100 SXM, 80 GB HBM3, per month
  • Same GPU, hourly on-demand$2,446market median × 730 h
  • You keep−49%8× node from $7,990 a month
  • Single-tenant bare metal, root access
  • Live in under 10 minutes after payment
  • No KYC · BTC, ETH, USDT
  • 99.9% uptime SLA · 24/7 engineers

Hardware

Which GPU for model training

Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.

GPUWhy for this workloadPer GPU, per monthAvailability
NVIDIA H100 SXM80 GB HBM3 · Hopper The default for anything up to 70B on an 8-GPU node: FP8 Transformer Engine, NVLink 4, 80 GB HBM3. $1,249 14 left
NVIDIA H200141 GB HBM3e · Hopper Same compute, 141 GB: larger micro-batches and longer sequences without more GPUs. $1,690 9 left
NVIDIA B200180 GB HBM3e · Blackwell Blackwell for the biggest runs: 180 GB and 8 TB/s per GPU, NVLink 5, FP4/FP8. $2,690 8 left
NVIDIA A10080 GB HBM2e · Ampere Fine-tunes and classic deep learning, where FP8 does not help and budget does. $649 21 left

Highlighted row: our default recommendation for model training. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.

Sizing

How much GPU memory does training need?

Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.

Full fine-tune, mixed precision~16 bytes per parameterWeights, gradients and Adam states: a 7B model needs ~112 GB — two 80 GB GPUs, or one B200.
LoRA16-bit weights + a few percent7B ≈ 16 GB, 13B ≈ 28 GB, 70B ≈ 140 GB — two H100s or one B300.
QLoRA, 4-bit base~0.6 bytes per parameter70B in about 40–48 GB: one A100 80 GB, or an L40S for small batches.
Activationsbatch × sequence lengthThe part that grows with context — and the reason H200 and B200 pay for themselves on long sequences.

From order to first run

  1. Pick the GPU and the termOne GPU for a fine-tune, an 8-GPU node for a pre-train; a 3-, 6- or 12-month term takes 5–15% off a run you know will last.
  2. Boot a ready imagePyTorch 2.7 with CUDA 12.6, the Hugging Face stack, JAX, or a bare Ubuntu with the driver — chosen at checkout, live in under 10 minutes.
  3. Stage data on local NVMe1–4 TB per server and 30 TB per node, as fast as the GPUs can read; attach block volumes for larger corpora.
  4. Checkpoint, monitor, renewCheckpoints stay on your NVMe, nvidia-smi and DCGM are yours, and the renewal settles from your balance so the run is never interrupted by a card.

In every server

  • Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
  • 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
  • Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
  • Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
  • Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
  • DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Read the documentation

Model training on a monthly GPU: questions answered

Is a monthly server cheaper than on-demand for training?

For a run that keeps the GPU busy, by a wide margin: an H100 SXM costs $1,249 a month here against roughly $2,446 for a full month at the market-median on-demand rate. The crossover is around 373 hours a month — below that, per-second billing wins, and we say so.

Can I train across more than one node?

Yes. Each node carries 3.2 Tb/s of InfiniBand for multi-node training with NCCL; for more than one node, open a sales ticket and we place them on the same fabric.

Do you sell spot or preemptible capacity?

No. Every server is dedicated and stays up for the whole term — the checkpoint you save at 03:00 is the one you resume from, not the one you lose.

Start model training tonight

Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.