H100 SXM vs L40S: Which GPU Should You Choose?
H100 SXM offers more device memory and NVLink than L40S; check whether either difference affects your workload before comparing rates.
Choose H100 SXM when its larger memory or a verified NVLink configuration is required. For a workload that fits L40S without those links, compare both GPUs using the current rates and measured completion cost.
Run a bounded test with the same model, precision, batch size and sequence length on each feasible option. Record peak memory, billed runtime and completed work; use the actual checkout rate to compare compute cost per completed job.
Today's prices
USD/hr · 1x GPU · prices grouped by provider and purchase model
| Provider | H100 SXM | L40S |
|---|---|---|
| Hyperstack | $3.20 · Not reported | Not offered |
| RunPod | $3.49 | $1.09 |
| DataCrunch | $3.85 | $1.59 |
| Vast.ai | $3.87 | $0.80 |
| Crusoe | $3.90 · Not reported | $1.50 · Not reported |
| Together AI | $3.99 · Not reported | Not offered |
| Paperspace | $5.95 · Not reported | Not offered |
| Lambda | Unavailable as of 2026-10-07 | Not offered |
Configuration prices used by workload rows
USD/hr · currently eligible on-demand configurations at the exact GPU count shown in the workload table
| Provider | GPU | Configuration | Price | Price basis | Availability |
|---|---|---|---|---|---|
| Lambda | H100 SXM | 2 GPUs | $8.38/hr | Published list rate | In stock |
Where the two parts differ
| Specification | H100 SXM | L40S |
|---|---|---|
| Memory | 80 GB HBM3 | 48 GB GDDR6 |
| Memory bandwidth | 3350 GB/s | 864 GB/s |
| Dense BF16 | 989 TFLOPS | 362 TFLOPS |
| Dense FP8 | 1979 TFLOPS | 733 TFLOPS |
| NVLink | 900 GB/s | No NVLink |
| TDP | 700 W | 350 W |
What your workload needs
Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows.
| Workload | VRAM needed | H100 SXM | L40S | Cheapest today |
|---|---|---|---|---|
| 7B Q4 inference | 4 GB | 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | L40S: 1 GPU at $0.8/hr |
| 13B Q4 inference | 8 GB | 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | L40S: 1 GPU at $0.8/hr |
| 70B QLoRA | 46 GB | 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | L40S: 1 GPU at $0.8/hr |
| 7B full fine-tune | 112 GB | 2-GPU configuration: $8.38/hr (In stock; observed 2026-10-07; price basis: Published list rate) | No eligible 3-GPU configuration | H100 SXM: 2 GPUs at $8.38/hr |
| 70B FP8 training | 154 GB | 2-GPU configuration: $8.38/hr (In stock; observed 2026-10-07; price basis: Published list rate) | No eligible 4-GPU configuration | H100 SXM: 2 GPUs at $8.38/hr |
What the numbers say
H100 SXM has a larger memory budget and NVLink support; L40S does not provide NVLink. These are configuration constraints before they are price arguments. If the workload fits one L40S, a multi-GPU interconnect advantage may be irrelevant; if it spans GPUs, verify how the offered machine connects them.
H100 SXM has an on-demand minimum of $3.20/hr and L40S has an on-demand minimum of $0.80/hr. The H100 SXM-to-L40S on-demand price ratio is 4. When both configurations meet the same requirements, compare the current rates rather than assuming that either GPU is always cheaper. Spot listings have different interruption terms and are not the basis of that ratio.
Measure a representative serving or training run on each feasible option. Record peak memory, billed time and completed work, including recovery overhead if you evaluate spot capacity. Do not move an NVLink-dependent job to L40S merely because a listing is available.
How to test the comparison
- Hold the workload constant
Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.
- Measure the result you need
For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.
- Price the complete run
Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.
Before you rent
Marlin matches your workload requirements to the lowest-priced suitable GPU across supported providers.
Memory fit and cost-per-job comparisons are models. Datasheet peaks do not establish application throughput; verify the listed configuration with your own workload. https://www.tensordock.com/host-pricing · https://www.coreweave.com/pricing · https://vast.ai/pricing · https://datacrunch.io/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413 · https://www.runpod.io/pricing · https://lambda.ai/service/gpu-cloud · https://www.hyperstack.cloud/gpu-pricing · https://www.together.ai/pricing · https://crusoe.ai/cloud/pricing
Marlin beta
Stop comparing. Start running.
Marlin matches your workload to the lowest-priced suitable GPU across supported providers.
Get started with Marlin