H200 SXM vs B200 SXM: Which GPU Fits Your Workload?
The right choice between H200 SXM and B200 SXM depends first on workload fit, then on current rates and measured job costs, using vendor specifications and observed provider listings as evidence.
When both configurations fit, compare the current on-demand rates and measured job costs; the current on-demand price ratio for H200 SXM relative to B200 SXM is 0.67. When the workload cannot fit in H200 SXM memory but can fit in B200 SXM memory, use B200 SXM. Choose B200 SXM for a topology that requires its higher NVLink bandwidth.
Before committing, run one bounded representative job with the same model, precision, batch size, sequence length, and output-quality target on each GPU, using the live table to verify the H200 SXM on-demand minimum of $3.99 USD/hr and the B200 SXM on-demand minimum of $6.00 USD/hr, then record peak memory, completed work, and billed runtime at each actual checkout rate.
Today's prices
USD/hr · 1x GPU · prices grouped by provider and purchase model
| Provider | H200 SXM | B200 SXM |
|---|---|---|
| Hyperstack | $3.99 · Not reported | $6.00 · Not reported |
| Crusoe | $4.29 · Not reported | Not offered |
| RunPod | $4.59 | $6.79 |
| DataCrunch | $5.02 | $7.27 |
| Together AI | $5.99 · Not reported | $8.19 · Not reported |
| Lambda | Not offered | Unavailable as of 2026-10-07 |
Where the two parts differ
| Specification | H200 SXM | B200 SXM |
|---|---|---|
| Memory | 141 GB HBM3e | 180 GB HBM3e |
| Memory bandwidth | 4800 GB/s | 8000 GB/s |
| Dense BF16 | 989.5 TFLOPS | 2250 TFLOPS |
| Dense FP8 | 1979 TFLOPS | 4500 TFLOPS |
| NVLink | 900 GB/s | 1800 GB/s |
| TDP | 700 W | 1000 W |
What your workload needs
Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows. 1 listed workload is omitted because neither card has a currently eligible on-demand configuration at the required GPU count.
| Workload | VRAM needed | H200 SXM | B200 SXM | Cheapest today |
|---|---|---|---|---|
| 7B Q4 inference | 4 GB | 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | H200 SXM: 1 GPU at $3.99/hr |
| 13B Q4 inference | 8 GB | 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | H200 SXM: 1 GPU at $3.99/hr |
| 7B full fine-tune | 112 GB | 1-GPU configuration: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | H200 SXM: 1 GPU at $3.99/hr |
| 70B FP8 training | 154 GB | No eligible 2-GPU configuration | 1-GPU configuration: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate) | B200 SXM: 1 GPU at $6/hr |
What the numbers say
H200 SXM has an on-demand minimum of $3.99 USD/hr, while B200 SXM has an on-demand minimum of $6.00 USD/hr. The on-demand price ratio for H200 SXM relative to B200 SXM is 0.67. When both configurations fit, compare the current on-demand rates and measured job costs before choosing.
B200 SXM anchors the screening model: the datasheet peak-performance ratio for H200 SXM relative to B200 SXM is 0.44, while the on-demand effective cost-per-job ratio for H200 SXM relative to B200 SXM is 1.51. These ratios use vendor peak specifications rather than measured throughput, so benchmark the same workload on both fitting configurations and record completed work and billed runtime.
B200 SXM provides a larger memory ceiling and higher NVLink bandwidth than H200 SXM. For FP8 or BF16 workloads within 180GB VRAM, B200 SXM covers the full stated envelope, while H200 SXM supports the same precision types only within its smaller memory ceiling. When the workload exceeds that smaller ceiling or the topology requires the higher NVLink bandwidth, choose B200 SXM; when both configurations fit, compare current on-demand rates and measured job costs.
How to test the comparison
- Hold the workload constant
Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.
- Measure the result you need
For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.
- Price the complete run
Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.
Before you rent
Marlin matches your workload requirements to the lowest-priced suitable GPU option across supported cloud providers.
Source coverage includes listable offers observed from supported providers; missing rows do not prove unavailability, and prices or capacity can change. Workload fit and cost ratios use stated memory assumptions, published rates, and datasheet peaks rather than measured completion times or job costs. https://www.hyperstack.cloud/gpu-pricing · https://www.together.ai/pricing · https://datacrunch.io/pricing · https://www.nvidia.com/en-us/data-center/h200/ · https://www.nvidia.com/en-us/data-center/hgx/ · https://www.nvidia.com/en-us/data-center/dgx-b200/ · https://www.runpod.io/pricing · https://lambda.ai/service/gpu-cloud · https://crusoe.ai/cloud/pricing · https://images.nvidia.com/aem-dam/Solutions/documents/HGX-B200-PCF-Summary.pdf
Marlin beta
Stop comparing. Start running.
Marlin matches your workload to the lowest-priced suitable GPU across supported providers.
Get started with Marlin