L40S vs RTX 4090: Which GPU Should You Choose?

L40S can keep larger workloads on one GPU than RTX 4090, while neither offers NVLink; measured fit should lead the shortlist.

· refreshed every 12 hours

Choose L40S when its larger memory keeps the job on one GPU. If the workload fits RTX 4090, compare current on-demand rates and measured completion cost. Neither GPU supplies NVLink.

Run a bounded test with the same model, precision, batch size and sequence length on each feasible option. Record peak memory, billed runtime and completed work; use the actual checkout rate to compare compute cost per completed job.

Does the job exceed RTX 4090 memory but fit one L40S? If yes, test L40S as a single-GPU configuration. If it fits either, benchmark both. If it exceeds both, verify another configuration rather than assuming multiple devices form one memory pool.
Which option meets your deadline and job budget in a representative test? Measure completion time first. The L40S-to-RTX 4090 modeled cost-per-job ratio is 1.01; it uses specification-based assumptions and does not establish real runtime or a deadline.
Can the exact configuration launch under the rental terms your job needs? Confirm the GPU count, topology and availability at checkout. Use spot only after testing checkpoint recovery; if the required configuration cannot launch, reassess fit before substituting another GPU.

Today's prices

USD/hr · 1x GPU · prices grouped by provider and purchase model

ProviderL40SRTX 4090
Vast.ai $0.80 $0.36
RunPod $1.09 $0.74
Crusoe $1.50 · Not reported Not offered
DataCrunch $1.59 Not offered

Where the two parts differ

SpecificationL40SRTX 4090
Memory48 GB GDDR624 GB GDDR6X
Memory bandwidth864 GB/s1008 GB/s
Dense BF16362 TFLOPS165 TFLOPS
Dense FP8733 TFLOPS330 TFLOPS
NVLinkNo NVLinkNo NVLink
TDP350 W450 W
Prices are refreshed every 12 hours. Marlin matches your requirements across supported providers. Try Marlin →
2.22×
L40S-to-RTX 4090 on-demand hourly price ratio. Below one favors the first GPU on hourly rate; above one favors the second.
1.01×
L40S-to-RTX 4090 modeled cost-per-job ratio. Below one favors the first under the stated model; validate with measured runtime.
$584/mo
Continuous-use monthly baseline for one L40S at its current on-demand minimum; scale the budget to your billed hours.

What your workload needs

Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows. 2 listed workloads are omitted because neither card has a currently eligible on-demand configuration at the required GPU count.

WorkloadVRAM neededL40SRTX 4090Cheapest today
7B Q4 inference 4 GB 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) 1-GPU configuration: $0.36/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) RTX 4090: 1 GPU at $0.36/hr
7B FP16 inference 17 GB 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) 1-GPU configuration: $0.36/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) RTX 4090: 1 GPU at $0.36/hr
13B LoRA fine-tune 36 GB 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) No eligible 2-GPU configuration L40S: 1 GPU at $0.8/hr

What the numbers say

L40S provides more device memory than RTX 4090, and neither provides NVLink. First measure whether model weights plus runtime state fit the smaller device. If they do not, the larger memory can avoid a split; if they exceed both, confirm an actual distributed configuration and its communication path.

L40S has an on-demand minimum of $0.80/hr and RTX 4090 has an on-demand minimum of $0.36/hr. The L40S-to-RTX 4090 on-demand price ratio is 2.22. When both configurations meet the same requirements, compare the current rates rather than assuming that either GPU is always cheaper. Spot listings have different interruption terms and are not the basis of that ratio.

For work that fits both, compare the same precision, batch size and output-quality target. A lower hourly rate can still cost more per completed run when billed runtime grows. Test serving latency and concurrency separately from an offline training throughput test.

How to test the comparison

  1. Hold the workload constant

    Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.

  2. Measure the result you need

    For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.

  3. Price the complete run

    Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.

Before you rent

Data egress charges sit outside the headline GPU rate; verify the provider's current per-GB outbound transfer fee on its live pricing page before estimating total cost.
A workload that exceeds single-GPU VRAM forces multi-GPU partitioning, increasing the rented GPU count and potentially adding communication overhead; budget against the full configuration, not one card.
Interruptible capacity can disappear mid-run, adding checkpoint recovery time and duplicate compute; use it only for workloads that can resume cleanly.
Billing granularity can make short experiments cost more than runtime alone suggests because minimum charges may round usage up; each provider's pricing page lists the applicable minimum charge.

Marlin matches your workload requirements to the lowest-priced suitable GPU across supported providers.

Memory fit and cost-per-job comparisons are models. Datasheet peaks do not establish application throughput; verify the listed configuration with your own workload. https://vast.ai/pricing · https://crusoe.ai/cloud/pricing · https://www.runpod.io/pricing · https://datacrunch.io/pricing · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413 · https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/

Marlin beta

Stop comparing. Start running.

Marlin matches your workload to the lowest-priced suitable GPU across supported providers.

Get started with Marlin