A100 80GB vs L40S: Which GPU Should You Choose?
A100 80GB and L40S trade larger memory and NVLink against native FP8 support; establish which constraint your job actually has.
Choose A100 80GB when the workload needs its larger memory or NVLink. Choose L40S when its memory is sufficient and your software needs native FP8; compare current job costs after establishing fit.
Run a bounded test with the same model, precision, batch size and sequence length on each feasible option. Record peak memory, billed runtime and completed work; use the actual checkout rate to compare compute cost per completed job.
Today's prices
USD/hr · 1x GPU · prices grouped by provider and purchase model
| Provider | A100 80GB | L40S |
|---|---|---|
| Vast.ai | $0.38 | $0.80 |
| Hyperstack | $1.35 · Not reported | Not offered |
| Jarvislabs | $1.49 · Not reported | Not offered |
| RunPod | $1.59 | $1.09 |
| DataCrunch | $1.85 | $1.59 |
| Crusoe | $2.00 · Not reported | $1.50 · Not reported |
Where the two parts differ
| Specification | A100 80GB | L40S |
|---|---|---|
| Memory | 80 GB HBM2e | 48 GB GDDR6 |
| Memory bandwidth | 2039 GB/s | 864 GB/s |
| Dense BF16 | 312 TFLOPS | 362 TFLOPS |
| Dense FP8 | FP8 unsupported | 733 TFLOPS |
| NVLink | 600 GB/s | No NVLink |
| TDP | 400 W | 350 W |
What your workload needs
Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows. 1 listed workload is omitted because neither card has a currently eligible on-demand configuration at the required GPU count.
| Workload | VRAM needed | A100 80GB | L40S | Cheapest today |
|---|---|---|---|---|
| 7B Q4 inference | 4 GB | 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | A100 80GB: 1 GPU at $0.38/hr |
| 13B Q4 inference | 8 GB | 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | A100 80GB: 1 GPU at $0.38/hr |
| 70B QLoRA | 46 GB | 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | 1-GPU configuration: $0.80/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) | A100 80GB: 1 GPU at $0.38/hr |
| 70B FP8 training | 154 GB | unsupported | No eligible 4-GPU configuration | No currently eligible on-demand configuration for L40S 4-GPU configuration |
What the numbers say
A100 80GB gives a larger single-GPU memory budget, while L40S offers native FP8 support. FP8 is a separate software-and-precision requirement: it does not make a workload fit when its total runtime memory exceeds the device capacity. Confirm both constraints before comparing prices.
A100 80GB has an on-demand minimum of $0.38/hr and L40S has an on-demand minimum of $0.80/hr. The A100 80GB-to-L40S on-demand price ratio is 0.47. When both configurations meet the same requirements, compare the current rates rather than assuming that either GPU is always cheaper. Spot listings have different interruption terms and are not the basis of that ratio.
For a BF16 job that fits either GPU, run the same model and batch on both. If the job requires native FP8, test the L40S software path; if it needs NVLink, verify an A100 configuration that exposes the required links. A GPU feature alone does not confirm the topology of a rental.
How to test the comparison
- Hold the workload constant
Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.
- Measure the result you need
For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.
- Price the complete run
Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.
Before you rent
Marlin matches your workload requirements to the lowest-priced suitable GPU across supported providers.
Memory fit and cost-per-job comparisons are models. Datasheet peaks do not establish application throughput; verify the listed configuration with your own workload. https://vast.ai/pricing · https://www.paperspace.com/pricing · https://datacrunch.io/pricing · https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413 · https://www.runpod.io/pricing · https://lambda.ai/service/gpu-cloud · https://www.hyperstack.cloud/gpu-pricing · https://jarvislabs.ai/pricing
Marlin beta
Stop comparing. Start running.
Marlin matches your workload to the lowest-priced suitable GPU across supported providers.
Get started with Marlin