L40S Rental Price: Cheapest Cloud per Hour (2026)

Choose L40S rental capacity for your workload using provider rates compiled from 5 sources and updated 2026-10-07.

· refreshed every 12 hours

For a one-GPU L40S workload that cannot tolerate interruption, choose Vast.ai at $0.80 per GPU-hour on on-demand; reported availability is In stock. If the same workload can restart after interruption, switch to DataCrunch at $0.80 per GPU-hour on spot; reported availability is In stock.

Recommendation source: This on-demand recommendation is a marketplace quote, not a provider-published list rate. Recheck it before booking.

Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.

Spot evidence: 1 observed spot configuration from 1 provider.

L40S is a GPU offered here as the priced rental unit 1x L40S. Every listed on-demand and spot offer uses the same GPU hardware, so a single-GPU workload must fit within that card's local memory; the purchase model changes rental terms, not GPU identity.

Before committing, run a 15-minute representative test, using $0.80 per GPU-hour as the observed on-demand reference rate; confirm the actual checkout rate, then record peak memory, completed work, billed runtime, and cost calculated from that runtime.

Do your model and serving workload fit within the VRAM capacity and GPU count shown in the adjacent workload table? If your workload matches FP8 or BF16 workloads within 48GB VRAM, apply the table's displayed formula to your exact model, precision, batch size, and sequence length, then validate peak memory on the displayed GPU count. If measured peak memory fits, confirm an observed offer for that exact configuration before considering a rental. If measured peak memory does not fit, increase the GPU count or change the model settings and repeat the test.
Can your run checkpoint its progress and restart after an interruption? If checkpointing is acceptable, confirm that an observed offer matches your exact GPU count and configuration. If it matches, consider DataCrunch at $0.80 per GPU-hour under spot; reported availability is In stock.
Do availability and price at checkout change your choice between Vast.ai at $0.80 per GPU-hour on on-demand (reported availability: In stock) and RunPod at $1.09 per GPU-hour on on-demand (reported availability: In stock)? If an observed offer matches your exact configuration, choose Vast.ai at $0.80 per GPU-hour on on-demand; reported availability is In stock. If both observed offers match, compare Vast.ai at $0.80 per GPU-hour on on-demand (reported availability: In stock) with RunPod at $1.09 per GPU-hour on on-demand (reported availability: In stock); the difference is $0.29 per GPU-hour and $2540.40 per year. If checkout availability or rates differ, choose among observed exact-configuration offers using current on-demand rates and measured job costs.

Current recommendation

For runs that cannot be interrupted, use Vast.ai at $0.80/hr (In stock).

The one observed spot offer is DataCrunch at $0.80/hr (In stock); use it only for checkpointable work.

The closest on-demand comparison is Vast.ai at $0.80/hr (In stock) versus RunPod at $1.09/hr (In stock): $0.29/hr, or $2540.40/year.

Today's prices

USD/hr · one observed offer per provider and pricing model

1 GPU

ProviderConfiguration $/hrPricingPrice basisAvailability
DataCrunch $0.80 spot Published list rate In stock
Vast.ai $0.80 on-demand Marketplace quote In stock
RunPod $1.09 on-demand Published list rate In stock
Crusoe $1.50 on-demand Published list rate Not reported
DataCrunch $1.59 on-demand Published list rate In stock
1.99x
Shortlist the observed same-GPU on-demand offers by current rate, then test the exact configuration and record completed work and billed runtime before choosing.
0.6%
Treat the spot versus on-demand rate gap as modeled, then compare it with measured checkpoint overhead, restart frequency, lost work, and billed runtime for your job.
$584/mo
Monthly on-demand spending scales with billed GPU-hours, so replace the full-time baseline with your measured workload runtime, GPU count, and actual checkout rate.

What an hour buys

Computed from the recommended on-demand offer: Vast.ai at $0.80/hr for 1 GPU.

$0.0167
Per GB of VRAM, hourly$0.8/GPU-hour divided by 48 GB of VRAM.
$0.0022
Per dense BF16 TFLOP, hourly$0.8/GPU-hour divided by 362 dense BF16 TFLOPS.
1080
GB/s of memory bandwidth per $/hour864 GB/s divided by $0.8/GPU-hour.

Hardware specifications

SpecificationL40S
Memory48 GB GDDR6
Memory bandwidth864 GB/s
Dense BF16362 TFLOPS
Dense FP8733 TFLOPS
NVLinkNo NVLink
TDP350 W

What your workload needs

VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 2 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.

Every row below: L40S: 1 GPU · Listed hourly cost: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote)

WorkloadVRAM needed
7B Q4 inference 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations)
13B Q4 inference 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations)
70B QLoRA 46 GB 70B × 0.5 B (4-bit base) × 1.3 (adapters+optimizer)

What the numbers say

L40S fit starts with the workload row that matches your model and precision: use its displayed memory formula and GPU count as a planning estimate, since each row has workload-specific assumptions. Validate that estimate with the same batch size, sequence length, and output-quality target you intend to run, and record peak memory before selecting a provider. If the workload exceeds the displayed configuration, change the model settings or GPU count and test again.

L40S has a 864GB/s memory-bandwidth specification for reasoning about bandwidth-bound decode, while its peak specifications of 362 BF16 TFLOPS and 733 FP8 TFLOPS frame compute-bound training or prefill. The 2.02x FP8-to-BF16 peak-compute ratio and FP8’s lower bytes per parameter indicate modeled potential, not measured fit or speed; both depend on model and software support. Benchmark the same model, precision, batch size, sequence length, and output-quality target, then record representative tokens or samples completed and billed runtime.

L40S fails the requirement if measured peak memory exceeds 48GB GDDR6 per GPU, if your software cannot execute the required BF16 or FP8 path, or if a multi-GPU run requires an interconnect that conflicts with the listed specification: No NVLink. Before choosing or rejecting it, test the exact model, precision, batch size, sequence length, and output-quality target; record peak memory and, for a multi-GPU configuration, measure communication overhead, completed work, and billed runtime.

Turn the workload estimate into a rental decision

  1. Measure the memory peak

    Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.

  2. Match the sold configuration

    Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.

  3. Bound the experiment

    Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.

How Marlin helps

Marlin matches users to the lowest-priced GPU option satisfying their workload requirements across supported CSP and GPU-cloud providers.

  • Supported-provider coverage: Marlin matches across every supported CSP and GPU-cloud provider, not only those with a listable price on this page.
  • Unified comparison: Marlin compares supported providers in one matching process, reducing the need to check prices manually across separate sites.
  • Requirement-aware choice: Marlin returns the lowest-priced GPU option that satisfies your workload requirements, rather than the cheapest row regardless of fit.

Manual price tracking means rechecking whether $0.80 per GPU-hour for on-demand still applies at checkout and repeating the comparison; Marlin handles that comparison when matching your workload to the lowest-priced suitable GPU option.

Before you rent

Persistent storage: The headline GPU rate does not show whether attached volumes or retained images cost extra, so verify any storage charge in the linked billing terms or checkout quote.
Checkpoint recovery: For spot runs, each interruption can add reload time and repeated compute, so measure lost work and billed runtime at your chosen checkpoint interval.
Configuration capacity: A collected listing does not guarantee the exact GPU count your workload needs at checkout, so verify the configuration and availability before scheduling the run.
Billing granularity: Per-second, per-minute, or per-hour rounding can make a short run cost more than its exact runtime at the quoted hourly rate, so confirm the billing increment on the provider's own billing page before committing.

Workload memory needs, minimum GPU counts, monthly and annual costs, and specification-derived performance ratios are modeled rather than measured. Listed rates were collected from provider pricing sources; re-check current rates through the supplied public provider pricing links or at provider checkout. Verify fixed hardware specifications against the supplied manufacturer links. https://vast.ai/pricing · https://www.runpod.io/pricing · https://datacrunch.io/pricing · https://crusoe.ai/cloud/pricing · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413

Marlin beta

Stop comparing. Start running.

Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable GPU option across supported providers.

Get started with Marlin