B200 SXM Rental Price: Cheapest Cloud per Hour (2026)
Decide where to rent B200 SXM using current offer data supported by 7 sources and updated 2026-10-07.
If your B200 SXM workload cannot restart after interruption, use the on-demand offer from Hyperstack at $6.00 per GPU-hour; availability is Not reported. If your B200 SXM workload can restart, use the spot alternative from DataCrunch at $3.63 per GPU-hour; availability is In stock.
Recommendation source: This on-demand recommendation uses a provider-published list rate; availability can still change.
Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.
Spot evidence: 1 observed spot configuration from 1 provider.
B200 SXM is a data-center GPU priced here as 1x B200 SXM. Because every offer uses the same hardware, workload memory fit does not change between on-demand and spot rentals; their rental terms change interruption and commitment conditions, not the GPU identity.
Run a bounded representative test with the planned model, precision, batch size, sequence length, and output-quality target, using $6.00 per GPU-hour as the on-demand reference rate; record peak memory, completed work, billed runtime, and the actual checkout rate before committing.
Current recommendation
For runs that cannot be interrupted, use Hyperstack at $6.00/hr (Not reported).
The one observed spot offer is DataCrunch at $3.63/hr (In stock); use it only for checkpointable work.
The closest on-demand comparison is Hyperstack at $6.00/hr (Not reported) versus RunPod at $6.79/hr (In stock): $0.79/hr, or $6920.40/year.
Today's prices
USD/hr · one observed offer per provider and pricing model
1 GPU
Every row below: Price basis: Published list rate
| Provider | Configuration $/hr | Pricing | Availability |
|---|---|---|---|
| DataCrunch | $3.63 | spot | In stock |
| Hyperstack | $6.00 | on-demand | Not reported |
| RunPod | $6.79 | on-demand | In stock |
| Lambda | $6.99 | on-demand | Unavailable as of 2026-10-07 |
| DataCrunch | $7.27 | on-demand | In stock |
| Together AI | $8.19 | on-demand | Not reported |
2 GPUs
| Provider | Configuration $/hr | Per GPU $/hr | Pricing | Price basis | Availability |
|---|---|---|---|---|---|
| Lambda | $13.78 | $6.89 | on-demand | Published list rate | Unavailable as of 2026-10-07 |
8 GPUs
| Provider | Configuration $/hr | Per GPU $/hr | Pricing | Price basis | Availability |
|---|---|---|---|---|---|
| Lambda | $53.52 | $6.69 | on-demand | Published list rate | Unavailable as of 2026-10-07 |
What an hour buys
Computed from the recommended on-demand offer: Hyperstack at $6.00/hr for 1 GPU.
Hardware specifications
| Specification | B200 SXM |
|---|---|
| Memory | 180 GB HBM3e |
| Memory bandwidth | 8000 GB/s |
| Dense BF16 | 2250 TFLOPS |
| Dense FP8 | 4500 TFLOPS |
| NVLink | 1800 GB/s |
| TDP | 1000 W |
What your workload needs
VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 1 listed workload is omitted because no currently eligible on-demand configuration exists at the required GPU count.
Every row below: B200 SXM: 1 GPU · Listed hourly cost: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate)
| Workload | VRAM needed |
|---|---|
| 7B Q4 inference | 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 13B Q4 inference | 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 70B FP8 training | 154 GB 70B × 2 B (FP8 mix + master/optimizer) × 1.1 (requires FP8 hardware) |
| 70B FP16 inference | 168 GB 70B × 2 B (FP16) × 1.2 (KV+activations) |
What the numbers say
B200 SXM fit depends on whether the planned model and serving settings stay within the capacity represented by the workload table. Use each row’s displayed formula and GPU count as an initial estimate because its assumptions are workload-specific, then test the intended model, precision, batch size, sequence length, and output-quality target while recording peak memory. Memory fit alone does not establish measured runtime or deadline performance.
B200 SXM datasheet peaks are specifications, not measured throughput: bandwidth-bound decode depends on memory movement, while compute-bound training or prefill depends on the precision path the software actually uses. The FP8-to-BF16 peak-throughput ratio of 2 and FP8’s lower bytes per parameter can change both fit and speed, but only when the software supports FP8 and the output-quality target is met. Interconnect capability is a separate concern for multi-GPU work. Benchmark the same model, precision, batch size, sequence length, and quality target to measure completed work and runtime.
B200 SXM should not be chosen from datasheet peaks alone. Reject the configuration if measured peak memory exceeds capacity, the software stack cannot use the intended precision, or the offered topology does not meet the workload’s interconnect needs. When another configuration also fits, compare current on-demand rates and measured job costs using the same model, precision, batch size, sequence length, and output-quality target. Rent only after confirming an observed offer for the exact configuration and its checkout terms.
Turn the workload estimate into a rental decision
- Measure the memory peak
Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.
- Match the sold configuration
Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.
- Bound the experiment
Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.
How Marlin helps
Marlin matches your workload requirements to the lowest-priced GPU option that satisfies them across supported CSP and GPU-cloud providers.
- Supported scope: Marlin matches across every supported CSP and GPU-cloud provider, including those without a listable offer on this page.
- One pass: Comparing supported providers in one matching process reduces the need to check each provider’s price separately.
- Qualified price: Marlin returns the cheapest option that satisfies your workload requirements, not the cheapest row regardless of fit.
This page implies manually checking provider rates against the on-demand reference of $6.00 per GPU-hour; Marlin handles that comparison when matching your workload to the lowest-priced suitable option.
Before you rent
Workload memory requirements, monthly and annualized costs, price ratios, and cost gaps are modeled from specifications and listed rates. Prices were collected from provider pricing sources, while memory, bandwidth, compute, and interconnect figures come from fixed manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://www.hyperstack.cloud/gpu-pricing · https://datacrunch.io/pricing · https://www.runpod.io/pricing · https://www.together.ai/pricing · https://www.nvidia.com/en-us/data-center/hgx/ · https://www.nvidia.com/en-us/data-center/dgx-b200/ · https://lambda.ai/service/gpu-cloud · https://images.nvidia.com/aem-dam/Solutions/documents/HGX-B200-PCF-Summary.pdf
Marlin beta
Stop comparing. Start running.
Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable option across supported providers.
Get started with Marlin