B200 SXM Rental Price: Cheapest Cloud per Hour (2026)

Decide where to rent B200 SXM using current offer data supported by 7 sources and updated 2026-10-07.

· refreshed every 12 hours

If your B200 SXM workload cannot restart after interruption, use the on-demand offer from Hyperstack at $6.00 per GPU-hour; availability is Not reported. If your B200 SXM workload can restart, use the spot alternative from DataCrunch at $3.63 per GPU-hour; availability is In stock.

Recommendation source: This on-demand recommendation uses a provider-published list rate; availability can still change.

Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.

Spot evidence: 1 observed spot configuration from 1 provider.

B200 SXM is a data-center GPU priced here as 1x B200 SXM. Because every offer uses the same hardware, workload memory fit does not change between on-demand and spot rentals; their rental terms change interruption and commitment conditions, not the GPU identity.

Run a bounded representative test with the planned model, precision, batch size, sequence length, and output-quality target, using $6.00 per GPU-hour as the on-demand reference rate; record peak memory, completed work, billed runtime, and the actual checkout rate before committing.

Does your model and serving workload fit B200 SXM at the VRAM requirement and GPU count shown in the adjacent workload table? If your workload meets FP8 or BF16 workloads within 180GB VRAM, use the table’s displayed formula and GPU count to estimate fit, then measure peak memory with the intended serving settings. If the measured workload fits, confirm an observed offer for that exact configuration before renting it. If it does not fit, revise the GPU count or workload settings and test again.
Can your run checkpoint frequently enough to restart after an interruption? If your run can checkpoint and restart, confirm an observed offer for 1x B200 SXM, then consider the spot option from DataCrunch at $3.63 per GPU-hour; its reported availability is In stock. If your run cannot restart, confirm an observed offer for the exact configuration before considering the on-demand option from Hyperstack at $6.00 per GPU-hour; its reported availability is Not reported.
After checking actual checkout availability and price, would your choice change between Hyperstack at $6.00 per GPU-hour for on-demand capacity (Not reported) and RunPod at $6.79 per GPU-hour for on-demand capacity (In stock)? If you confirm an observed offer for 1x B200 SXM at checkout, use the supplied on-demand recommendation: Hyperstack at $6.00 per GPU-hour, with reported availability of Not reported. When both configurations fit, compare the current on-demand rates and measured job costs: Hyperstack at $6.00 per GPU-hour (Not reported) versus RunPod at $6.79 per GPU-hour (In stock). The displayed rates differ by $0.79 per GPU-hour, with a modeled annual difference of $6920.40 per year.

Current recommendation

For runs that cannot be interrupted, use Hyperstack at $6.00/hr (Not reported).

The one observed spot offer is DataCrunch at $3.63/hr (In stock); use it only for checkpointable work.

The closest on-demand comparison is Hyperstack at $6.00/hr (Not reported) versus RunPod at $6.79/hr (In stock): $0.79/hr, or $6920.40/year.

Today's prices

USD/hr · one observed offer per provider and pricing model

1 GPU

Every row below: Price basis: Published list rate

ProviderConfiguration $/hrPricingAvailability
DataCrunch $3.63 spot In stock
Hyperstack $6.00 on-demand Not reported
RunPod $6.79 on-demand In stock
Lambda $6.99 on-demand Unavailable as of 2026-10-07
DataCrunch $7.27 on-demand In stock
Together AI $8.19 on-demand Not reported

2 GPUs

ProviderConfiguration $/hrPer GPU $/hrPricingPrice basisAvailability
Lambda $13.78 $6.89 on-demand Published list rate Unavailable as of 2026-10-07

8 GPUs

ProviderConfiguration $/hrPer GPU $/hrPricingPrice basisAvailability
Lambda $53.52 $6.69 on-demand Published list rate Unavailable as of 2026-10-07
1.36x
Use the on-demand provider spread to shortlist current offers, then compare any modeled spot savings with measured restart and checkpoint costs before choosing a purchase model.
39.4%
Treat the displayed spot-to-on-demand cost gap as modeled, then compare it with measured checkpoint and restart costs before choosing either purchase model.
$4380/mo
Use the monthly on-demand baseline as a budgeting model rather than a bill forecast, then compare the modeled gap with measured restart and checkpoint costs before choosing between on-demand and spot capacity.

What an hour buys

Computed from the recommended on-demand offer: Hyperstack at $6.00/hr for 1 GPU.

$0.0333
Per GB of VRAM, hourly$6/GPU-hour divided by 180 GB of VRAM.
$0.0027
Per dense BF16 TFLOP, hourly$6/GPU-hour divided by 2250 dense BF16 TFLOPS.
1333
GB/s of memory bandwidth per $/hour8000 GB/s divided by $6/GPU-hour.

Hardware specifications

SpecificationB200 SXM
Memory180 GB HBM3e
Memory bandwidth8000 GB/s
Dense BF162250 TFLOPS
Dense FP84500 TFLOPS
NVLink1800 GB/s
TDP1000 W

What your workload needs

VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 1 listed workload is omitted because no currently eligible on-demand configuration exists at the required GPU count.

Every row below: B200 SXM: 1 GPU · Listed hourly cost: $6.00/hr (Not reported; observed 2026-10-07; price basis: Published list rate)

WorkloadVRAM needed
7B Q4 inference 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations)
13B Q4 inference 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations)
70B FP8 training 154 GB 70B × 2 B (FP8 mix + master/optimizer) × 1.1 (requires FP8 hardware)
70B FP16 inference 168 GB 70B × 2 B (FP16) × 1.2 (KV+activations)

What the numbers say

B200 SXM fit depends on whether the planned model and serving settings stay within the capacity represented by the workload table. Use each row’s displayed formula and GPU count as an initial estimate because its assumptions are workload-specific, then test the intended model, precision, batch size, sequence length, and output-quality target while recording peak memory. Memory fit alone does not establish measured runtime or deadline performance.

B200 SXM datasheet peaks are specifications, not measured throughput: bandwidth-bound decode depends on memory movement, while compute-bound training or prefill depends on the precision path the software actually uses. The FP8-to-BF16 peak-throughput ratio of 2 and FP8’s lower bytes per parameter can change both fit and speed, but only when the software supports FP8 and the output-quality target is met. Interconnect capability is a separate concern for multi-GPU work. Benchmark the same model, precision, batch size, sequence length, and quality target to measure completed work and runtime.

B200 SXM should not be chosen from datasheet peaks alone. Reject the configuration if measured peak memory exceeds capacity, the software stack cannot use the intended precision, or the offered topology does not meet the workload’s interconnect needs. When another configuration also fits, compare current on-demand rates and measured job costs using the same model, precision, batch size, sequence length, and output-quality target. Rent only after confirming an observed offer for the exact configuration and its checkout terms.

Turn the workload estimate into a rental decision

  1. Measure the memory peak

    Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.

  2. Match the sold configuration

    Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.

  3. Bound the experiment

    Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.

How Marlin helps

Marlin matches your workload requirements to the lowest-priced GPU option that satisfies them across supported CSP and GPU-cloud providers.

  • Supported scope: Marlin matches across every supported CSP and GPU-cloud provider, including those without a listable offer on this page.
  • One pass: Comparing supported providers in one matching process reduces the need to check each provider’s price separately.
  • Qualified price: Marlin returns the cheapest option that satisfies your workload requirements, not the cheapest row regardless of fit.

This page implies manually checking provider rates against the on-demand reference of $6.00 per GPU-hour; Marlin handles that comparison when matching your workload to the lowest-priced suitable option.

Before you rent

Storage: Verify whether persistent volumes are billed separately from the GPU rate in the linked billing terms or checkout quote.
Restarts: spot interruptions can add checkpoint reload and recompute time, so measure completed work against billed runtime.
Capacity: A collected listing is not a reservation, so confirm checkout availability for 1x B200 SXM before planning the run.
Billing: Per-second, per-minute, or per-hour rounding can change the effective cost of short runs, so confirm the billing increment on the provider’s own billing page before committing.

Workload memory requirements, monthly and annualized costs, price ratios, and cost gaps are modeled from specifications and listed rates. Prices were collected from provider pricing sources, while memory, bandwidth, compute, and interconnect figures come from fixed manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://www.hyperstack.cloud/gpu-pricing · https://datacrunch.io/pricing · https://www.runpod.io/pricing · https://www.together.ai/pricing · https://www.nvidia.com/en-us/data-center/hgx/ · https://www.nvidia.com/en-us/data-center/dgx-b200/ · https://lambda.ai/service/gpu-cloud · https://images.nvidia.com/aem-dam/Solutions/documents/HGX-B200-PCF-Summary.pdf

Marlin beta

Stop comparing. Start running.

Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable option across supported providers.

Get started with Marlin