Launch

Launch pricing: up to 74% below on-demand, fixed for your term — ends Oct 31, 2026 · 42d left Launch pricing ends Oct 31 · 42d left

See pricing

Ada Lovelace · data center GPU

Rent a dedicated NVIDIA L40S server for $429 a month

The L40S is the universal data center GPU of the Ada generation: 48 GB of GDDR6 with ECC, fourth-generation tensor cores with FP8, third-generation RT cores for rendering, and three NVENC/NVDEC engines for video. It runs LLM inference, image and video generation, virtual workstations and transcoding on the same card — which is why it is our recommended GPU for always-on inference. Dedicated bare metal, 12 vCPU, 96 GB of RAM and 1 TB of NVMe included.

  • Per GPU, per month$429$0.59/h effective · fixed for your term
  • Same GPU, hourly on-demand$1,124market median $1.54/h × 730 h
  • You keep−62%$695 a month, every month
  • In stock · 17 left
  • Live in under 10 minutes after payment
  • No KYC · BTC, ETH, USDT
  • 99.9% uptime SLA · 24/7 engineers

Specification

NVIDIA L40S server specifications

One dedicated machine, sized around the GPU. Nothing is shared, nothing is metered, and every line below is included in the monthly price.

GPU memory48 GB GDDR6864 GB/s memory bandwidth
Tensor compute362 TFLOPS FP16Ada Lovelace architecture
InterconnectPCIe 5.0Single-GPU servers, full x16 lanes
Dedicated host12 vCPU · 96 GB RAM · 1 TB NVMeBare metal, single tenant, root access
Network10 Gbps uplink20 TB outbound included · IPv4 + /64 IPv6 · DDoS protection
ImagesUbuntu 24.04 · CUDA 12.6PyTorch, TensorFlow, vLLM, Triton, K3s at checkout
Regions5 Tier III data centersAshburn, Dallas, Amsterdam, Frankfurt, Stockholm
Availability17 leftBatch of 40 · counters updated 18 September 2026

Why teams rent the L40S here

  • 48 GB with FP8. A 30B-class model at FP8, or a 70B model at 4-bit, on one GPU — with the Transformer Engine formats the H100 uses.
  • RT cores and video engines. Ray-traced rendering, Omniverse and three NVENC/NVDEC pairs: the same card serves graphics and AI.
  • Inference-class price. About a third of an H100 SXM per month, for workloads that do not need HBM bandwidth.
  • One flat invoice. $429 a month covers the GPU, the host, the NVMe, 20 TB of transfer and 24/7 support — the same number every month, fixed for your term.

Below about 279 hours of use a month, an hourly provider is the cheaper way to run this GPU. Above it — and a machine that trains, serves or renders is above it — the monthly rate wins, and the gap is the $695 shown above.

A GPU server tray pulled out of its rack in the data center
Every L40S is delivered as a whole machine — one tenant per server, yours for the term.

Pricing

L40S price per month, by term

Per GPU, in USD, excluding VAT. Longer terms cost less; every term includes the same dedicated host, network and support.

TermPer GPU, per monthEffective hourlyvs on-demandOrder
MonthlyRolling month-to-month. Cancel with 30 days' notice. $429 $0.59/h −62% Deploy
3 months5% off the monthly rate for a 3-month term. $408 $0.56/h −64% Deploy
6 months10% off the monthly rate for a 6-month term. $386 $0.53/h −66% Deploy
12 months15% off the monthly rate for a 12-month term. $365 $0.50/h −68% Deploy

Launch pricing: these rates apply to orders confirmed before Oct 31, 2026 and stay fixed for your whole term. From Nov 1, new L40S orders are billed at the list price of $499 a month. On-demand comparison: market-median published hourly rate for the same GPU ($1.54/h) over 730 hours. Block storage $25/TB/month and extra IPv4 addresses $4/month are optional add-ons.

Market comparison

L40S rental price compared

Every provider in our audit that publishes an on-demand rate for the L40S, read on 19 September 2026 from their own pricing page and multiplied by 730 hours — what the same GPU costs kept for a month.

ProviderPublished rate, L40SOne month, 24/7vs our $429/mo
GPU Cloud HQDedicated, billed monthly $429 per month, flat $429 Our reference
DigitalOcean (Paperspace)On-demand, hourly $1.57/GPU/h $1,146 −63%
NebiusOn-demand, hourly from $1.55/h on demand · from $0.74 preemptible $1,132 −62%
RunPodOn-demand, hourly $1.09/h Secure Cloud · $0.79 Community $796 −46%
Market median30–40 providers, GetDeploying index $1.54/h on demand $1,124 −62%

Rates are each provider's standard on-demand tier for the L40S — not spot, preemptible or community hardware — as printed on their pricing page on 19 September 2026; each row links to the full comparison with its source. Where a provider is cheaper than us for a full month, the row says so. Method and caveats: GPU cloud alternatives.

Workloads

What people run on an L40S

Inference, video & graphics — and the three jobs below are where a dedicated L40S at a flat monthly price earns its keep.

  • Always-on LLM inference

    vLLM or TensorRT-LLM endpoints for 7B–30B models at FP8, chat and RAG backends, embedding and reranking services.

  • Image and video generation

    SDXL, FLUX and video-diffusion pipelines with room for several models resident, plus hardware encode for the output.

  • Virtual workstations and transcoding

    Remote 3D workstations, streaming and live transcoding on NVENC — work a training GPU is overqualified for.

Is the L40S the right GPU for you?

Choose the L40S for inference and graphics that fit in 48 GB and run around the clock. If a 70B model at FP8 is the target, the H200 or a pair of H100s is the honest answer; if the job is small-model inference or transcoding only, the L4 at a fraction of the price is enough.

Recommended for inference serving — the sizing guide there shows what fits in 48 GB GDDR6.

L40S rental: questions answered

How much does it cost to rent an L40S server?

$429 a month, flat, for a dedicated NVIDIA L40S server with 12 vCPU · 96 GB RAM · 1 TB NVMe and a 10 Gbps uplink — $0.59 per hour effective over 730 hours. A 3-, 6- or 12-month term takes 5%, 10% or 15% off: $408, $386 or $365 a month. The market-median on-demand rate for the same GPU is $1.54 per hour, about $1,124 for a full month.

Is the L40S in stock right now?

Yes. 17 of the current batch of 40 were unallocated at the last stock update (18 September 2026). Provisioning is automated: your server is imaged, secured with your SSH key and handed over in under 10 minutes after your payment confirms.

What is included with a L40S server?

Everything on the pricing table: the GPU, 12 vCPU · 96 GB RAM · 1 TB NVMe, a 10 Gbps uplink with 20 TB of outbound transfer, a dedicated public IPv4 and a /64 IPv6 block, DDoS protection, root access and Ubuntu 24.04 with the NVIDIA driver, CUDA 12 and Docker — or PyTorch, vLLM, Kubernetes images at checkout. Hardware replacement within 4 hours and a 99.9% uptime SLA are part of the contract.

Can the L40S run LLM inference?

Yes — it is our recommended GPU for it. With 48 GB and FP8 support it serves models up to roughly 30B parameters at FP8, or 70B at 4-bit, on one card. It lacks the HBM bandwidth of an H100, so per-token latency on very large batches is higher, but for a steady production endpoint the flat monthly price is hard to argue with.

Can I rent several L40Ss?

Yes: order 1, 2, 4 or 8 GPUs at once, subject to the units left in the batch, on the same monthly terms.

Do I need KYC or a card to rent it?

No. An email address is all we ask for. You fund a prepaid balance in Bitcoin, Ether or USDT (ERC-20 or TRC-20); the first invoice settles from it and monthly renewals charge to it automatically. No ID document, no card, no bank transfer.

Deploy an L40S in under 10 minutes

$429 a month, fixed for your term. Fund your balance in BTC, ETH or USDT and the server is imaged, secured with your key and handed over automatically.