L40S Rental Price: Cheapest Cloud per Hour (2026)
Choose L40S rental capacity for your workload using provider rates compiled from 5 sources and updated 2026-10-07.
For a one-GPU L40S workload that cannot tolerate interruption, choose Vast.ai at $0.80 per GPU-hour on on-demand; reported availability is In stock. If the same workload can restart after interruption, switch to DataCrunch at $0.80 per GPU-hour on spot; reported availability is In stock.
Recommendation source: This on-demand recommendation is a marketplace quote, not a provider-published list rate. Recheck it before booking.
Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.
Spot evidence: 1 observed spot configuration from 1 provider.
L40S is a GPU offered here as the priced rental unit 1x L40S. Every listed on-demand and spot offer uses the same GPU hardware, so a single-GPU workload must fit within that card's local memory; the purchase model changes rental terms, not GPU identity.
Before committing, run a 15-minute representative test, using $0.80 per GPU-hour as the observed on-demand reference rate; confirm the actual checkout rate, then record peak memory, completed work, billed runtime, and cost calculated from that runtime.
Current recommendation
For runs that cannot be interrupted, use Vast.ai at $0.80/hr (In stock).
The one observed spot offer is DataCrunch at $0.80/hr (In stock); use it only for checkpointable work.
The closest on-demand comparison is Vast.ai at $0.80/hr (In stock) versus RunPod at $1.09/hr (In stock): $0.29/hr, or $2540.40/year.
Today's prices
USD/hr · one observed offer per provider and pricing model
1 GPU
| Provider | Configuration $/hr | Pricing | Price basis | Availability |
|---|---|---|---|---|
| DataCrunch | $0.80 | spot | Published list rate | In stock |
| Vast.ai | $0.80 | on-demand | Marketplace quote | In stock |
| RunPod | $1.09 | on-demand | Published list rate | In stock |
| Crusoe | $1.50 | on-demand | Published list rate | Not reported |
| DataCrunch | $1.59 | on-demand | Published list rate | In stock |
What an hour buys
Computed from the recommended on-demand offer: Vast.ai at $0.80/hr for 1 GPU.
Hardware specifications
| Specification | L40S |
|---|---|
| Memory | 48 GB GDDR6 |
| Memory bandwidth | 864 GB/s |
| Dense BF16 | 362 TFLOPS |
| Dense FP8 | 733 TFLOPS |
| NVLink | No NVLink |
| TDP | 350 W |
What your workload needs
VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 2 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.
Every row below: L40S: 1 GPU · Listed hourly cost: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote)
| Workload | VRAM needed |
|---|---|
| 7B Q4 inference | 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 13B Q4 inference | 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 70B QLoRA | 46 GB 70B × 0.5 B (4-bit base) × 1.3 (adapters+optimizer) |
What the numbers say
L40S fit starts with the workload row that matches your model and precision: use its displayed memory formula and GPU count as a planning estimate, since each row has workload-specific assumptions. Validate that estimate with the same batch size, sequence length, and output-quality target you intend to run, and record peak memory before selecting a provider. If the workload exceeds the displayed configuration, change the model settings or GPU count and test again.
L40S has a 864GB/s memory-bandwidth specification for reasoning about bandwidth-bound decode, while its peak specifications of 362 BF16 TFLOPS and 733 FP8 TFLOPS frame compute-bound training or prefill. The 2.02x FP8-to-BF16 peak-compute ratio and FP8’s lower bytes per parameter indicate modeled potential, not measured fit or speed; both depend on model and software support. Benchmark the same model, precision, batch size, sequence length, and output-quality target, then record representative tokens or samples completed and billed runtime.
L40S fails the requirement if measured peak memory exceeds 48GB GDDR6 per GPU, if your software cannot execute the required BF16 or FP8 path, or if a multi-GPU run requires an interconnect that conflicts with the listed specification: No NVLink. Before choosing or rejecting it, test the exact model, precision, batch size, sequence length, and output-quality target; record peak memory and, for a multi-GPU configuration, measure communication overhead, completed work, and billed runtime.
Turn the workload estimate into a rental decision
- Measure the memory peak
Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.
- Match the sold configuration
Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.
- Bound the experiment
Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.
How Marlin helps
Marlin matches users to the lowest-priced GPU option satisfying their workload requirements across supported CSP and GPU-cloud providers.
- Supported-provider coverage: Marlin matches across every supported CSP and GPU-cloud provider, not only those with a listable price on this page.
- Unified comparison: Marlin compares supported providers in one matching process, reducing the need to check prices manually across separate sites.
- Requirement-aware choice: Marlin returns the lowest-priced GPU option that satisfies your workload requirements, rather than the cheapest row regardless of fit.
Manual price tracking means rechecking whether $0.80 per GPU-hour for on-demand still applies at checkout and repeating the comparison; Marlin handles that comparison when matching your workload to the lowest-priced suitable GPU option.
Before you rent
Workload memory needs, minimum GPU counts, monthly and annual costs, and specification-derived performance ratios are modeled rather than measured. Listed rates were collected from provider pricing sources; re-check current rates through the supplied public provider pricing links or at provider checkout. Verify fixed hardware specifications against the supplied manufacturer links. https://vast.ai/pricing · https://www.runpod.io/pricing · https://datacrunch.io/pricing · https://crusoe.ai/cloud/pricing · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413
Marlin beta
Stop comparing. Start running.
Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable GPU option across supported providers.
Get started with Marlin