Launch

Launch pricing: up to 74% below on-demand, fixed for your term — ends Oct 31, 2026 · 42d left Launch pricing ends Oct 31 · 42d left

See pricing

Solutions · Fine-tuning & RLHF

GPU servers for fine-tuning and RLHF, at a flat monthly price

Fine-tuning queues never empty: one experiment leads to the next, evaluation runs in between, and the GPU that was “for one job” is busy for months. A dedicated 80 GB server at a flat rate turns that queue into a fixed line on the budget.

  • Recommended GPU$649NVIDIA A100, 80 GB HBM2e, per month
  • Same GPU, hourly on-demand$1,285market median × 730 h
  • You keep−49%$636 a month, every month
  • Single-tenant bare metal, root access
  • Live in under 10 minutes after payment
  • No KYC · BTC, ETH, USDT
  • 99.9% uptime SLA · 24/7 engineers

Hardware

Which GPU for fine-tuning & rlhf

Four picks from the catalogue, each a dedicated server with host, NVMe, network and support included. Click through for specs, price by term and the market comparison.

GPUWhy for this workloadPer GPU, per monthAvailability
NVIDIA A10080 GB HBM2e · Ampere The value pick: 80 GB for LoRA, QLoRA and full fine-tunes up to 13B, at about half the price of an H100. $649 21 left
NVIDIA H100 SXM80 GB HBM3 · Hopper Twice the throughput and FP8, for full fine-tunes and DPO/PPO loops that must finish this week. $1,249 14 left
NVIDIA H200141 GB HBM3e · Hopper 141 GB: full fine-tunes of 30B-class models, and RLHF with policy and reference models on one card. $1,690 9 left
NVIDIA RTX 409024 GB GDDR6X · Ada Lovelace LoRA and QLoRA on 7B–13B models when 24 GB is enough and the budget is small. $189 19 left

Highlighted row: our default recommendation for fine-tuning & rlhf. Prices per GPU per month in USD, excl. VAT; 3-, 6- and 12-month terms take 5–15% off. Stock counters updated 18 September 2026.

Sizing

Which GPU for which fine-tune?

Rules of thumb we use when a customer asks which GPU to order. Memory decides; everything else is speed.

QLoRA, 4-bit base~0.6 bytes per parameter70B in about 40–48 GB: one A100 80 GB, or an L40S for small batches.
LoRA, 16-bit base2 bytes per parameter + adapters13B ≈ 28 GB on a 4090 or L40S; 70B ≈ 140 GB on two H100s or one B300.
Full fine-tune~16 bytes per parameter7B ≈ 112 GB (two 80 GB GPUs or a B200); 70B ≈ 1.1 TB (an 8× H200, B200 or B300 node).
DPO and RLHFpolicy + reference (+ reward)DPO on 7B–13B fits one A100 or H100; PPO with three resident models is H200 or multi-GPU territory.

From order to first run

  1. Choose memory, not just speedThe table above is the whole decision: the adapter method sets the memory, the memory sets the GPU, and the GPU sets the flat monthly price.
  2. Start from the Hugging Face imageTransformers, PEFT, Accelerate and bitsandbytes on PyTorch 2.7 — nothing to install before the first run.
  3. Keep datasets and adapters localNVMe scratch sized to the GPU, block volumes for shared corpora, and 20 TB of transfer a month to push results wherever they go.
  4. Evaluate on the same boxvLLM is one image away: serve the adapter you just trained from the same server for evaluation and A/B tests.

In every server

  • Single-tenant bare metal. Your GPU, CPU cores, RAM and NVMe are yours alone. No noisy neighbours, no oversubscription.
  • 10–25 Gbps uplink. Dedicated bandwidth per server with 20 TB of outbound transfer included every month. No egress fees.
  • Local NVMe storage. Fast scratch space for datasets and checkpoints, sized to the GPU. Expand with attached block volumes.
  • Ready-to-train images. Ubuntu with NVIDIA drivers, CUDA 12, cuDNN and Docker — or PyTorch, TensorFlow, vLLM, Kubernetes and Windows Server images at checkout.
  • Public IPv4 & IPv6. A dedicated IPv4 address and a /64 IPv6 block on every server, with reverse DNS on request.
  • DDoS protection. Always-on volumetric mitigation at the network edge, included at no charge.
Read the documentation

Fine-tuning & RLHF on a monthly GPU: questions answered

Which GPU do I need to fine-tune a 70B model?

QLoRA fits a 70B model on a single A100 80 GB or, for small batches, an L40S. LoRA at 16-bit needs about 140 GB — two H100 SXMs or one B300. A full fine-tune needs about 1.1 TB: an 8-GPU H200, B200 or B300 node, or optimizer offloading on an H100 node.

Can I run RLHF or DPO on one server?

DPO on a 7B–13B model fits one A100 or H100. PPO-style RLHF, which keeps policy, reference and reward models resident, is where the H200 (141 GB) or a multi-GPU order pays off.

Do you charge for moving datasets in and out?

No. 20 TB of outbound transfer a month is included on every server and inbound traffic is unmetered. Local NVMe is included; block storage is $25 per TB a month.

Start fine-tuning & rlhf tonight

Every GPU above is in stock. Fund your balance, pick the server, and be at a root prompt in under 10 minutes.