H200 SXM Rental Price: Cheapest Cloud per Hour (2026)

Choose an H200 SXM rental by memory fit, interruption tolerance and the price of an actual launchable configuration. Current rates and reported availability appear below.

· refreshed every 12 hours

For H200 SXM workloads that cannot restart, choose Hyperstack at $3.99 with availability listed as Not reported; if the workload can restart, switch to the spot alternative from DataCrunch at $2.51 with availability listed as In stock.

Recommendation source: This on-demand recommendation uses a provider-published list rate; availability can still change.

Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.

Spot evidence: 1 observed spot configuration from 1 provider.

H200 SXM is the exact GPU variant compared here. A per-GPU comparison does not establish that a larger node can be rented one GPU at a time.

Before committing, run a bounded test at the confirmed checkout rate. Record peak memory, billed runtime and completed work. Include checkpoint recovery in the test if you intend to use spot.

Does your model plus serving overhead require more than 80GB but fit within 141GB? Verify that the workload meets FP8 or BF16 workloads within 141GB VRAM, then use the workload table to test fit: 141GB raises the fit boundary versus 80GB, but KV cache and activation needs remain workload-dependent.
Can the run checkpoint and restart after an interruption? If checkpointing is acceptable, use the spot option from DataCrunch at $2.51 with availability listed as In stock.
Do the actual price and availability facts change your choice between Hyperstack at $3.99 (Not reported) and Crusoe at $4.29 (Not reported)? Use the on-demand recommendation, Hyperstack at $3.99 with availability listed as Not reported. The lowest offer is Hyperstack at $3.99 (Not reported), while the runner-up is Crusoe at $4.29 (Not reported), a difference of $0.30 per hour and $2628.00 per year.

Current recommendation

For runs that cannot be interrupted, use Hyperstack at $3.99/hr (Not reported).

The one observed spot offer is DataCrunch at $2.51/hr (In stock); use it only for checkpointable work.

The closest on-demand comparison is Hyperstack at $3.99/hr (Not reported) versus Crusoe at $4.29/hr (Not reported): $0.30/hr, or $2628.00/year.

Today's prices

USD/hr · one observed offer per provider and pricing model

1 GPU

Every row below: Price basis: Published list rate

ProviderConfiguration $/hrPricingAvailability
DataCrunch $2.51 spot In stock
Hyperstack $3.99 on-demand Not reported
Crusoe $4.29 on-demand Not reported
RunPod $4.59 on-demand In stock
DataCrunch $5.02 on-demand In stock
Together AI $5.99 on-demand Not reported
1.5x
The on-demand price spread makes provider comparison material, so use $3.99 as the low-end baseline and weigh availability before choosing.
37%
For restartable work, compare the modeled spot cost with on-demand and test whether recovery overhead changes the result.
$2913/mo
For continuous use, compare the full on-demand provider range before committing because the yearly gap measured on this page reaches $17520.00.

What an hour buys

Computed from the recommended on-demand offer: Hyperstack at $3.99/hr for 1 GPU.

$0.0283
Per GB of VRAM, hourly$3.99/GPU-hour divided by 141 GB of VRAM.
$0.0040
Per dense BF16 TFLOP, hourly$3.99/GPU-hour divided by 989.5 dense BF16 TFLOPS.
1203
GB/s of memory bandwidth per $/hour4800 GB/s divided by $3.99/GPU-hour.

Hardware specifications

SpecificationH200 SXM
Memory141 GB HBM3e
Memory bandwidth4800 GB/s
Dense BF16989.5 TFLOPS
Dense FP81979 TFLOPS
NVLink900 GB/s
TDP700 W

What your workload needs

VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 2 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.

Every row below: H200 SXM: 1 GPU · Listed hourly cost: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate)

WorkloadVRAM needed
7B Q4 inference 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations)
13B Q4 inference 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations)
7B full fine-tune 112 GB 7B × 16 B (FP16 weights+grads+Adam fp32 states)

What the numbers say

H200 SXM provides a 141GB HBM3e capacity boundary versus H100 SXM at 80GB HBM3, so use the existing workload table to decide whether the added memory changes model fit. The rows use different workload-specific assumptions. Measure KV cache, activations and runtime overhead at your intended batch size and sequence length. If the workload fits the smaller boundary, H100 SXM at $3.20 may be the cheaper alternative.

H200 SXM has a specification-level memory-bandwidth bound of 4800GB/s versus H100 SXM at 3350GB/s, which can change a bandwidth-bound decode path, while compute-bound training or prefill may see unchanged compute speed because their specified BF16 figures are 989.5 and 989 TFLOPS and their FP8 figures are 1979 and 1979 TFLOPS. FP8’s 2x throughput relationship and lower bytes per parameter change the same modeled fit-and-speed choice. The 900GB/s and 900GB/s interconnect specifications are separate from this single-GPU decision. HBM generation changes capacity and bandwidth, not compute, and none of these specifications represents measured workload throughput.

H200 SXM provides 141GB of HBM3e, 4800GB/s of memory bandwidth, and 900GB/s NVLink, while L40S has 48GB of GDDR6, 864GB/s, and No NVLink, and RTX 4090 has 24GB of GDDR6X, 1008GB/s, and No NVLink. H200 NVL matches the 141GB HBM3e capacity, 4800GB/s bandwidth, and 900GB/s NVLink, but its 835.5 BF16 TFLOPS, 1670.5 FP8 TFLOPS, and 600W differ from this GPU’s 989.5 BF16 TFLOPS, 1979 FP8 TFLOPS, and 700W, so it is not a direct substitute. When the model fits the smaller device and training is compute-bound, benchmark both and compare current checkout rates before paying for additional memory.

Turn the workload estimate into a rental decision

  1. Measure the memory peak

    Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.

  2. Match the sold configuration

    Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.

  3. Bound the experiment

    Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.

How Marlin helps

Marlin matches workload requirements to the lowest-priced suitable GPU option across its supported providers.

  • Search breadth: Marlin matches across every supported provider, including providers beyond those priced on this page.
  • Unified comparison: Marlin compares supported providers in one matching process, reducing the need to check prices manually across separate sites.
  • Qualified lowest price: Marlin returns the cheapest option that satisfies the reader's workload requirements, not simply the cheapest row regardless of fit.

This page implies manually checking provider prices and validating the $3.99 rate before committing; Marlin handles the comparison when matching a workload to the lowest-priced suitable option.

Before you rent

Storage: Check the linked storage or billing terms and the checkout quote to establish whether volumes and snapshots are charged separately from GPU time.
Checkpoint recovery: For spot runs, checkpoint frequency and restart time can add paid recomputation, so test recovery with the intended workload before estimating total cost.
Launchable configuration: Verify the exact GPU count, topology and availability before scheduling the run. A memory estimate does not establish that the configuration can be rented.
Billing granularity: Per-second, per-minute, or per-hour rounding can change the effective cost of short runs, so confirm the billing increment on the provider's own billing page before committing.

Workload memory and cost projections use the stated assumptions; validate them with measured memory and billed runtime. Prices were collected from provider pricing sources, while hardware figures come from fixed manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://www.hyperstack.cloud/gpu-pricing · https://datacrunch.io/pricing · https://crusoe.ai/cloud/pricing · https://www.runpod.io/pricing · https://www.together.ai/pricing · https://www.nvidia.com/en-us/data-center/h200/ · https://vast.ai/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413 · https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/

Marlin beta

Stop comparing. Start running.

Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable option across supported providers.

Get started with Marlin