H200 SXM vs B200 SXM: Which GPU Fits Your Workload?

The right choice between H200 SXM and B200 SXM depends first on workload fit, then on current rates and measured job costs, using vendor specifications and observed provider listings as evidence.

· refreshed every 12 hours

When both configurations fit, compare the current on-demand rates and measured job costs; the current on-demand price ratio for H200 SXM relative to B200 SXM is 0.67. When the workload cannot fit in H200 SXM memory but can fit in B200 SXM memory, use B200 SXM. Choose B200 SXM for a topology that requires its higher NVLink bandwidth.

Before committing, run one bounded representative job with the same model, precision, batch size, sequence length, and output-quality target on each GPU, using the live table to verify the H200 SXM on-demand minimum of $3.99 USD/hr and the B200 SXM on-demand minimum of $6.00 USD/hr, then record peak memory, completed work, and billed runtime at each actual checkout rate.

How much GPU memory does your workload require at peak? H200 SXM supports FP8 or BF16 workloads within 141GB VRAM; B200 SXM supports FP8 or BF16 workloads within 180GB VRAM. When both configurations fit, you should compare current on-demand rates and measured job costs only for exact configurations with observed offers. When H200 SXM does not fit but B200 SXM does, you should rent B200 SXM only after confirming an observed offer for the exact configuration.
Is finishing sooner worth the measured per-job cost difference? The on-demand effective cost-per-job ratio for H200 SXM relative to B200 SXM is 1.51, with H200 SXM as the numerator, and is a screening model based on datasheet BF16 peaks rather than measured runtime. When both configurations fit, you should benchmark the same model, precision, batch size, sequence length, and output-quality target on each against the same deadline. If one configuration meets the deadline at a lower measured job cost, you should rent it only after confirming an observed offer for that exact configuration.
Is the exact configuration available for your required window? If availability is uncertain, you should first verify an observed offer for the exact configuration and required window in the provider’s checkout flow.

Today's prices

USD/hr · 1x GPU · prices grouped by provider and purchase model

ProviderH200 SXMB200 SXM
Hyperstack $3.99 · Not reported $6.00 · Not reported
Crusoe $4.29 · Not reported Not offered
RunPod $4.59 $6.79
DataCrunch $5.02 $7.27
Together AI $5.99 · Not reported $8.19 · Not reported
Lambda Not offered Unavailable as of 2026-10-07

Where the two parts differ

SpecificationH200 SXMB200 SXM
Memory141 GB HBM3e180 GB HBM3e
Memory bandwidth4800 GB/s8000 GB/s
Dense BF16989.5 TFLOPS2250 TFLOPS
Dense FP81979 TFLOPS4500 TFLOPS
NVLink900 GB/s1800 GB/s
TDP700 W1000 W
Prices are refreshed every 12 hours. Marlin matches your requirements across supported providers. Try Marlin →
0.67×
Use the on-demand price ratio to screen hourly spend only when both configurations fit, then decide with measured job costs.
1.51×
Use the on-demand effective cost model to prioritize benchmarks; measured runtime and completed work determine actual job cost.
$2913/mo
Use this on-demand monthly run rate to budget sustained usage, then reconcile it with the actual checkout rate and billed runtime.

What your workload needs

Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows. 1 listed workload is omitted because neither card has a currently eligible on-demand configuration at the required GPU count.

WorkloadVRAM neededH200 SXMB200 SXMCheapest today
7B Q4 inference 4 GB 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H200 SXM: 1 GPU at $3.99/hr
13B Q4 inference 8 GB 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H200 SXM: 1 GPU at $3.99/hr
7B full fine-tune 112 GB 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H200 SXM: 1 GPU at $3.99/hr
70B FP8 training 154 GB No eligible 2-GPU configuration 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) B200 SXM: 1 GPU at $6/hr

What the numbers say

H200 SXM has an on-demand minimum of $3.99 USD/hr, while B200 SXM has an on-demand minimum of $6.00 USD/hr. The on-demand price ratio for H200 SXM relative to B200 SXM is 0.67. When both configurations fit, compare the current on-demand rates and measured job costs before choosing.

B200 SXM anchors the screening model: the datasheet peak-performance ratio for H200 SXM relative to B200 SXM is 0.44, while the on-demand effective cost-per-job ratio for H200 SXM relative to B200 SXM is 1.51. These ratios use vendor peak specifications rather than measured throughput, so benchmark the same workload on both fitting configurations and record completed work and billed runtime.

B200 SXM provides a larger memory ceiling and higher NVLink bandwidth than H200 SXM. For FP8 or BF16 workloads within 180GB VRAM, B200 SXM covers the full stated envelope, while H200 SXM supports the same precision types only within its smaller memory ceiling. When the workload exceeds that smaller ceiling or the topology requires the higher NVLink bandwidth, choose B200 SXM; when both configurations fit, compare current on-demand rates and measured job costs.

How to test the comparison

  1. Hold the workload constant

    Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.

  2. Measure the result you need

    For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.

  3. Price the complete run

    Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.

Before you rent

Data-egress charges are separate from headline GPU rates; verify the provider’s current per-GB egress terms and estimate them from the bytes your job will transfer.
A workload that exceeds single-GPU memory may require multiple GPUs, adding GPU-hours and communication overhead to the total job cost.
Interruptible spot capacity can end before a job completes, and the resulting lost work or restart time can change the total cost.
Billing granularity can make a short experiment cost more than its runtime implies; each provider’s pricing page states the minimum charge needed to budget the smallest test.

Marlin matches your workload requirements to the lowest-priced suitable GPU option across supported cloud providers.

Source coverage includes listable offers observed from supported providers; missing rows do not prove unavailability, and prices or capacity can change. Workload fit and cost ratios use stated memory assumptions, published rates, and datasheet peaks rather than measured completion times or job costs. https://www.hyperstack.cloud/gpu-pricing · https://www.together.ai/pricing · https://datacrunch.io/pricing · https://www.nvidia.com/en-us/data-center/h200/ · https://www.nvidia.com/en-us/data-center/hgx/ · https://www.nvidia.com/en-us/data-center/dgx-b200/ · https://www.runpod.io/pricing · https://lambda.ai/service/gpu-cloud · https://crusoe.ai/cloud/pricing · https://images.nvidia.com/aem-dam/Solutions/documents/HGX-B200-PCF-Summary.pdf

Marlin beta

Stop comparing. Start running.

Marlin matches your workload to the lowest-priced suitable GPU across supported providers.

Get started with Marlin