Solutions · Model training
GPU servers for model training, billed by the month
A training run is the definition of a machine that stays busy: weeks of wall-clock time at full utilisation, checkpoints every few hours, and a bill that hourly clouds price for the idle capacity they keep on standby. A monthly server or node removes the meter — and the idle.
- Recommended GPU$1,249NVIDIA H100 SXM, 80 GB HBM3, per month
- Same GPU, hourly on-demand$2,446market median × 730 h
- You keep−49%8× node from $7,990 a month
Hardware
Which GPU for model training
Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.
| GPU | Why for this workload | Per GPU, per month | Availability |
|---|---|---|---|
| NVIDIA H100 SXM80 GB HBM3 · Hopper | The default for anything up to 70B on an 8-GPU node: FP8 Transformer Engine, NVLink 4, 80 GB HBM3. | $1,249 | 14 left |
| NVIDIA H200141 GB HBM3e · Hopper | Same compute, 141 GB: larger micro-batches and longer sequences without more GPUs. | $1,690 | 9 left |
| NVIDIA B200180 GB HBM3e · Blackwell | Blackwell for the biggest runs: 180 GB and 8 TB/s per GPU, NVLink 5, FP4/FP8. | $2,690 | 8 left |
| NVIDIA A10080 GB HBM2e · Ampere | Fine-tunes and classic deep learning, where FP8 does not help and budget does. | $649 | 21 left |
Highlighted row: our default recommendation for model training. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.
Sizing
How much GPU memory does training need?
Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.
From order to first run
- Pick the GPU and the termOne GPU for a fine-tune, an 8-GPU node for a pre-train; a 3-, 6- or 12-month term takes 5–15% off a run you know will last.
- Boot a ready imagePyTorch 2.7 with CUDA 12.6, the Hugging Face stack, JAX, or a bare Ubuntu with the driver — chosen at checkout, live in under 10 minutes.
- Stage data on local NVMe1–4 TB per server and 30 TB per node, as fast as the GPUs can read; attach block volumes for larger corpora.
- Checkpoint, monitor, renewCheckpoints stay on your NVMe, nvidia-smi and DCGM are yours, and the renewal settles from your balance so the run is never interrupted by a card.
In every server
- Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
- 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
- Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
- Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
- Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
- DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Model training on a monthly GPU: questions answered
Is a monthly server cheaper than on-demand for training?
For a run that keeps the GPU busy, by a wide margin: an H100 SXM costs $1,249 a month here against roughly $2,446 for a full month at the market-median on-demand rate. The crossover is around 373 hours a month — below that, per-second billing wins, and we say so.
Can I train across more than one node?
Yes. Each node carries 3.2 Tb/s of InfiniBand for multi-node training with NCCL; for more than one node, open a sales ticket and we place them on the same fabric.
Do you sell spot or preemptible capacity?
No. Every server is dedicated and stays up for the whole term — the checkpoint you save at 03:00 is the one you resume from, not the one you lose.
Start model training tonight
Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.