H100 SXM vs A100 80GB: Which GPU Should You Choose?

H100 SXM adds native FP8 to the comparison with A100 80GB; BF16 jobs need a measured benefit to justify a different rental cost.

· refreshed every 12 hours

Choose H100 SXM when your software requires its native FP8 support. For BF16 work that fits either GPU, benchmark both and compare current total job costs before choosing.

Run a bounded test with the same model, precision, batch size and sequence length on each feasible option. Record peak memory, billed runtime and completed work; use the actual checkout rate to compare compute cost per completed job.

Does your validated software path require native FP8? For native FP8, shortlist H100 SXM and verify output quality. For BF16 work that fits either GPU, keep both for a representative runtime and cost test.
Which option meets your deadline and job budget in a representative test? Measure completion time first. The H100 SXM-to-A100 80GB modeled cost-per-job ratio is 2.66; it uses specification-based assumptions and does not establish real runtime or a deadline.
Can the exact configuration launch under the rental terms your job needs? Confirm the GPU count, topology and availability at checkout. Use spot only after testing checkpoint recovery; if the required configuration cannot launch, reassess fit before substituting another GPU.

Today's prices

USD/hr · 1x GPU · prices grouped by provider and purchase model

ProviderH100 SXMA100 80GB
Hyperstack $3.20 · Not reported $1.35 · Not reported
RunPod $3.49 $1.59
DataCrunch $3.85 $1.85
Vast.ai $3.87 $0.38
Crusoe $3.90 · Not reported $2.00 · Not reported
Together AI $3.99 · Not reported Not offered
Paperspace $5.95 · Not reported Not offered
Jarvislabs Not offered $1.49 · Not reported
Lambda Unavailable as of 2026-10-07 Not offered

Configuration prices used by workload rows

USD/hr · currently eligible on-demand configurations at the exact GPU count shown in the workload table

ProviderGPUConfigurationPricePrice basisAvailability
CoreWeave H100 SXM 8 GPUs $49.24/hr Published list rate Not reported
Lambda H100 SXM 2 GPUs $8.38/hr Published list rate In stock

see all H100 prices →

Where the two parts differ

SpecificationH100 SXMA100 80GB
Memory80 GB HBM380 GB HBM2e
Memory bandwidth3350 GB/s2039 GB/s
Dense BF16989 TFLOPS312 TFLOPS
Dense FP81979 TFLOPSFP8 unsupported
NVLink900 GB/s600 GB/s
TDP700 W400 W
Prices are refreshed every 12 hours. Marlin matches your requirements across supported providers. Try Marlin →
8.42×
H100 SXM-to-A100 80GB on-demand hourly price ratio. Below one favors the first GPU on hourly rate; above one favors the second.
2.66×
H100 SXM-to-A100 80GB modeled cost-per-job ratio. Below one favors the first under the stated model; validate with measured runtime.
$2336/mo
Continuous-use monthly baseline for one H100 SXM at its current on-demand minimum; scale the budget to your billed hours.

What your workload needs

Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows.

WorkloadVRAM neededH100 SXMA100 80GBCheapest today
7B Q4 inference 4 GB 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) A100 80GB: 1 GPU at $0.38/hr
13B Q4 inference 8 GB 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) A100 80GB: 1 GPU at $0.38/hr
70B QLoRA 46 GB 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $0.38/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) A100 80GB: 1 GPU at $0.38/hr
70B FP8 training 154 GB 2-GPU configuration: $8.38/hr (In stock; observed 2026-10-07; price basis: Published list rate) unsupported H100 SXM: 2 GPUs at $8.38/hr
70B FP16 inference 168 GB 8-GPU configuration: $49.24/hr (Not reported; observed 2026-10-07; price basis: Published list rate) No eligible 8-GPU configuration H100 SXM: 8 GPUs at $49.24/hr

What the numbers say

H100 SXM and A100 80GB share a recorded memory capacity, but the A100 lacks native FP8 support. FP8 therefore changes the shortlist only when the model, kernels and accuracy target support that precision. A BF16 workload still needs its own runtime measurement.

H100 SXM has an on-demand minimum of $3.20/hr and A100 80GB has an on-demand minimum of $0.38/hr. The H100 SXM-to-A100 80GB on-demand price ratio is 8.42. When both configurations meet the same requirements, compare the current rates rather than assuming that either GPU is always cheaper. Spot listings have different interruption terms and are not the basis of that ratio.

For the BF16 comparison, hold model, batch size and sequence length constant and record billed runtime. For an FP8 evaluation, also check output quality against your reference. Choose H100 for an established capability need or measured benefit, rather than treating a datasheet peak as a promised speedup.

How to test the comparison

  1. Hold the workload constant

    Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.

  2. Measure the result you need

    For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.

  3. Price the complete run

    Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.

Before you rent

Data egress fees can raise the final bill; verify the provider's current per-GB transfer charge on its billing page.
A model that does not fit on one GPU forces a multi-GPU rental, multiplying billed GPU-hours for every run.
Interruptible rentals can stop mid-run, and infrequent checkpoints may turn discounted capacity into repeated compute.
Short experiments can cost more when usage rounds up to a billing increment; each provider's pricing page lists the minimum charge needed for an accurate estimate.

Marlin matches your workload requirements to the lowest-priced suitable GPU across supported providers.

Memory fit and cost-per-job comparisons are models. Datasheet peaks do not establish application throughput; verify the listed configuration with your own workload. https://www.tensordock.com/host-pricing · https://www.coreweave.com/pricing · https://vast.ai/pricing · https://www.paperspace.com/pricing · https://datacrunch.io/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://www.nvidia.com/content/dam/en-zz/Solutions/Data-Center/a100/pdf/nvidia-a100-datasheet-us-nvidia-1758950-r4-web.pdf · https://www.runpod.io/pricing · https://lambda.ai/service/gpu-cloud · https://www.hyperstack.cloud/gpu-pricing · https://www.together.ai/pricing · https://crusoe.ai/cloud/pricing · https://jarvislabs.ai/pricing

Marlin beta

Stop comparing. Start running.

Marlin matches your workload to the lowest-priced suitable GPU across supported providers.

Get started with Marlin