H200 SXM Rental Price: Cheapest Cloud per Hour (2026)
Choose an H200 SXM rental by memory fit, interruption tolerance and the price of an actual launchable configuration. Current rates and reported availability appear below.
For H200 SXM workloads that cannot restart, choose Hyperstack at $3.99 with availability listed as Not reported; if the workload can restart, switch to the spot alternative from DataCrunch at $2.51 with availability listed as In stock.
Recommendation source: This on-demand recommendation uses a provider-published list rate; availability can still change.
Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.
Spot evidence: 1 observed spot configuration from 1 provider.
H200 SXM is the exact GPU variant compared here. A per-GPU comparison does not establish that a larger node can be rented one GPU at a time.
Before committing, run a bounded test at the confirmed checkout rate. Record peak memory, billed runtime and completed work. Include checkpoint recovery in the test if you intend to use spot.
Current recommendation
For runs that cannot be interrupted, use Hyperstack at $3.99/hr (Not reported).
The one observed spot offer is DataCrunch at $2.51/hr (In stock); use it only for checkpointable work.
The closest on-demand comparison is Hyperstack at $3.99/hr (Not reported) versus Crusoe at $4.29/hr (Not reported): $0.30/hr, or $2628.00/year.
Today's prices
USD/hr · one observed offer per provider and pricing model
1 GPU
Every row below: Price basis: Published list rate
| Provider | Configuration $/hr | Pricing | Availability |
|---|---|---|---|
| DataCrunch | $2.51 | spot | In stock |
| Hyperstack | $3.99 | on-demand | Not reported |
| Crusoe | $4.29 | on-demand | Not reported |
| RunPod | $4.59 | on-demand | In stock |
| DataCrunch | $5.02 | on-demand | In stock |
| Together AI | $5.99 | on-demand | Not reported |
What an hour buys
Computed from the recommended on-demand offer: Hyperstack at $3.99/hr for 1 GPU.
Hardware specifications
| Specification | H200 SXM |
|---|---|
| Memory | 141 GB HBM3e |
| Memory bandwidth | 4800 GB/s |
| Dense BF16 | 989.5 TFLOPS |
| Dense FP8 | 1979 TFLOPS |
| NVLink | 900 GB/s |
| TDP | 700 W |
What your workload needs
VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 2 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.
Every row below: H200 SXM: 1 GPU · Listed hourly cost: $3.99/hr (Not reported; observed 2026-10-07; price basis: Published list rate)
| Workload | VRAM needed |
|---|---|
| 7B Q4 inference | 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 13B Q4 inference | 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 7B full fine-tune | 112 GB 7B × 16 B (FP16 weights+grads+Adam fp32 states) |
What the numbers say
H200 SXM provides a 141GB HBM3e capacity boundary versus H100 SXM at 80GB HBM3, so use the existing workload table to decide whether the added memory changes model fit. The rows use different workload-specific assumptions. Measure KV cache, activations and runtime overhead at your intended batch size and sequence length. If the workload fits the smaller boundary, H100 SXM at $3.20 may be the cheaper alternative.
H200 SXM has a specification-level memory-bandwidth bound of 4800GB/s versus H100 SXM at 3350GB/s, which can change a bandwidth-bound decode path, while compute-bound training or prefill may see unchanged compute speed because their specified BF16 figures are 989.5 and 989 TFLOPS and their FP8 figures are 1979 and 1979 TFLOPS. FP8’s 2x throughput relationship and lower bytes per parameter change the same modeled fit-and-speed choice. The 900GB/s and 900GB/s interconnect specifications are separate from this single-GPU decision. HBM generation changes capacity and bandwidth, not compute, and none of these specifications represents measured workload throughput.
H200 SXM provides 141GB of HBM3e, 4800GB/s of memory bandwidth, and 900GB/s NVLink, while L40S has 48GB of GDDR6, 864GB/s, and No NVLink, and RTX 4090 has 24GB of GDDR6X, 1008GB/s, and No NVLink. H200 NVL matches the 141GB HBM3e capacity, 4800GB/s bandwidth, and 900GB/s NVLink, but its 835.5 BF16 TFLOPS, 1670.5 FP8 TFLOPS, and 600W differ from this GPU’s 989.5 BF16 TFLOPS, 1979 FP8 TFLOPS, and 700W, so it is not a direct substitute. When the model fits the smaller device and training is compute-bound, benchmark both and compare current checkout rates before paying for additional memory.
Turn the workload estimate into a rental decision
- Measure the memory peak
Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.
- Match the sold configuration
Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.
- Bound the experiment
Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.
How Marlin helps
Marlin matches workload requirements to the lowest-priced suitable GPU option across its supported providers.
- Search breadth: Marlin matches across every supported provider, including providers beyond those priced on this page.
- Unified comparison: Marlin compares supported providers in one matching process, reducing the need to check prices manually across separate sites.
- Qualified lowest price: Marlin returns the cheapest option that satisfies the reader's workload requirements, not simply the cheapest row regardless of fit.
This page implies manually checking provider prices and validating the $3.99 rate before committing; Marlin handles the comparison when matching a workload to the lowest-priced suitable option.
Before you rent
Workload memory and cost projections use the stated assumptions; validate them with measured memory and billed runtime. Prices were collected from provider pricing sources, while hardware figures come from fixed manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://www.hyperstack.cloud/gpu-pricing · https://datacrunch.io/pricing · https://crusoe.ai/cloud/pricing · https://www.runpod.io/pricing · https://www.together.ai/pricing · https://www.nvidia.com/en-us/data-center/h200/ · https://vast.ai/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://resources.nvidia.com/en-us-l40s/l40s-datasheet-28413 · https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
Marlin beta
Stop comparing. Start running.
Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable option across supported providers.
Get started with Marlin