Solutions · Fine-tuning & RLHF
GPU servers for fine-tuning and RLHF, at a flat monthly price
Fine-tuning queues never empty: one experiment leads to the next, evaluation runs in between, and the GPU that was “for one job” is busy for months. A dedicated 80 GB server at a flat rate turns that queue into a fixed line on the budget.
- Recommended GPU$649NVIDIA A100, 80 GB HBM2e, per month
- Same GPU, hourly on-demand$1,285market median × 730 h
- You keep−49%$636 a month, every month
Hardware
Which GPU for fine-tuning & rlhf
Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.
| GPU | Why for this workload | Per GPU, per month | Availability |
|---|---|---|---|
| NVIDIA A10080 GB HBM2e · Ampere | The value pick: 80 GB for LoRA, QLoRA and full fine-tunes up to 13B, at about half the price of an H100. | $649 | 21 left |
| NVIDIA H100 SXM80 GB HBM3 · Hopper | Twice the throughput and FP8, for full fine-tunes and DPO/PPO loops that must finish this week. | $1,249 | 14 left |
| NVIDIA H200141 GB HBM3e · Hopper | 141 GB: full fine-tunes of 30B-class models, and RLHF with policy and reference models on one card. | $1,690 | 9 left |
| NVIDIA RTX 409024 GB GDDR6X · Ada Lovelace | LoRA and QLoRA on 7B–13B models when 24 GB is enough and the budget is small. | $189 | 19 left |
Highlighted row: our default recommendation for fine-tuning & rlhf. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.
Sizing
Which GPU for which fine-tune?
Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.
From order to first run
- Choose memory, not just speedThe table above is the whole decision: the adapter method sets the memory, the memory sets the GPU, and the GPU sets the flat monthly price.
- Start from the Hugging Face imageTransformers, PEFT, Accelerate and bitsandbytes on PyTorch 2.7 — nothing to install before the first run.
- Keep datasets and adapters localNVMe scratch sized to the GPU, block volumes for shared corpora, and 20 TB of transfer a month to push results wherever they go.
- Evaluate on the same boxvLLM is one image away: serve the adapter you just trained from the same server for evaluation and A/B tests.
In every server
- Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
- 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
- Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
- Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
- Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
- DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Fine-tuning & RLHF on a monthly GPU: questions answered
Which GPU do I need to fine-tune a 70B model?
QLoRA fits a 70B model on a single A100 80 GB or, for small batches, an L40S. LoRA at 16-bit needs about 140 GB — two H100 SXMs or one B300. A full fine-tune needs about 1.1 TB: an 8-GPU H200, B200 or B300 node, or optimizer offloading on an H100 node.
Can I run RLHF or DPO on one server?
DPO on a 7B–13B model fits one A100 or H100. PPO-style RLHF, which keeps policy, reference and reward models resident, is where the H200 (141 GB) or a multi-GPU order pays off.
Do you charge for moving datasets in and out?
No. 20 TB of outbound transfer a month is included on every server and inbound traffic is unmetered. Local NVMe is included; block storage is $25 per TB a month.
Start fine-tuning & rlhf tonight
Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.