H100 SXM Rental Price: Cheapest Cloud per Hour (2026)

Choose an H100 SXM rental by memory fit, interruption tolerance and the price of an actual launchable configuration. Current rates and reported availability appear below.

· refreshed every 12 hours

For H100 SXM, choose Hyperstack at $3.20 (Not reported); if the workload can restart, use the spot alternative from DataCrunch at $1.93 (In stock), but keep the on-demand choice when interruptions would prevent recovery.

Recommendation source: This on-demand recommendation uses a provider-published list rate; availability can still change.

Spot recommendation source: This spot recommendation uses a provider-published list rate; availability can still change.

Spot evidence: 2 observed spot configurations from 1 provider; this is a configuration count, not a count of independent provider quotes.

H100 SXM is the exact GPU variant compared here. A per-GPU comparison does not establish that a larger node can be rented one GPU at a time.

Before committing, run a bounded test at the confirmed checkout rate. Record peak memory, billed runtime and completed work. Include checkpoint recovery in the test if you intend to use spot.

Does your model and serving workload fit the VRAM requirement and GPU count shown in the adjacent workload table? Confirm that your use case matches FP8 or BF16 workloads within 80GB VRAM, then check the applicable workload formula against your actual memory needs and validate the displayed GPU count with a representative run.
Can your run checkpoint reliably and restart after an interruption? If checkpointing is acceptable, use the spot option from DataCrunch at $1.93 (In stock); otherwise, stay with on-demand capacity.
Do the actual availability and checkout price change your choice between Hyperstack at $3.20 (Not reported) and RunPod at $3.49 (In stock)? The on-demand recommendation is Hyperstack at $3.20 (Not reported). Compare Hyperstack at $3.20 (Not reported) with RunPod at $3.49 (In stock): the difference is $0.29 per hour and $2540.40 per year.

Current recommendation

For runs that cannot be interrupted, use Hyperstack at $3.20/hr (Not reported).

2 spot offers are observed; the lowest is DataCrunch at $1.93/hr (In stock); use it only for checkpointable work.

The closest on-demand comparison is Hyperstack at $3.20/hr (Not reported) versus RunPod at $3.49/hr (In stock): $0.29/hr, or $2540.40/year.

Today's prices

USD/hr · one observed offer per provider and pricing model

1 GPU

ProviderConfiguration $/hrPricingPrice basisAvailability
DataCrunch $1.93 spot Published list rate In stock
Hyperstack $3.20 on-demand Published list rate Not reported
RunPod $3.49 on-demand Published list rate In stock
DataCrunch $3.85 on-demand Published list rate In stock
Vast.ai $3.87 on-demand Marketplace quote In stock
Crusoe $3.90 on-demand Published list rate Not reported
Together AI $3.99 on-demand Published list rate Not reported
Lambda $4.29 on-demand Published list rate Unavailable as of 2026-10-07
Paperspace $5.95 on-demand Published list rate Not reported

2 GPUs

ProviderConfiguration $/hrPer GPU $/hrPricingPrice basisRegionAvailability
Lambda $8.38 $4.19 on-demand Published list rate us-southeast-1 In stock

4 GPUs

ProviderConfiguration $/hrPer GPU $/hrPricingPrice basisAvailability
Lambda $16.36 $4.09 on-demand Published list rate Unavailable as of 2026-10-07

8 GPUs

Every row below: Pricing: on-demand · Price basis: Published list rate

ProviderConfiguration $/hrPer GPU $/hrAvailability
Lambda $31.92 $3.99 Unavailable as of 2026-10-07
CoreWeave $49.24 $6.16 Not reported
1.86x
The spread in on-demand prices makes provider selection a material cost decision, so verify availability and terms near the $3.20 low end before committing.
39.8%
For restartable work, compare the modeled spot cost with on-demand and test whether recovery overhead changes the result.
$2336/mo
Continuous use turns the lowest on-demand rate into a recurring monthly commitment, so budget against that baseline and verify availability before committing; the measured gap between the lowest and highest tracked rates reaches $24090.00 per year.

What an hour buys

Computed from the recommended on-demand offer: Hyperstack at $3.20/hr for 1 GPU.

$0.0400
Per GB of VRAM, hourly$3.2/GPU-hour divided by 80 GB of VRAM.
$0.0032
Per dense BF16 TFLOP, hourly$3.2/GPU-hour divided by 989 dense BF16 TFLOPS.
1047
GB/s of memory bandwidth per $/hour3350 GB/s divided by $3.2/GPU-hour.

Hardware specifications

SpecificationH100 SXM
Memory80 GB HBM3
Memory bandwidth3350 GB/s
Dense BF16989 TFLOPS
Dense FP81979 TFLOPS
NVLink900 GB/s
TDP700 W

What your workload needs

VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count.

WorkloadVRAM neededH100 SXMListed hourly cost
7B Q4 inference 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations) 1 GPU $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate)
13B Q4 inference 8 GB 13B × 0.5 B (4-bit) × 1.2 (KV+activations) 1 GPU $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate)
70B QLoRA 46 GB 70B × 0.5 B (4-bit base) × 1.3 (adapters+optimizer) 1 GPU $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate)
70B FP8 training 154 GB 70B × 2 B (FP8 mix + master/optimizer) × 1.1 (requires FP8 hardware) 2 GPUs $8.38/hr (In stock; observed 2026-10-07; price basis: Published list rate)
70B FP16 inference 168 GB 70B × 2 B (FP16) × 1.2 (KV+activations) 8 GPUs $49.24/hr (Not reported; observed 2026-10-07; price basis: Published list rate)

What the numbers say

H100 SXM should be selected only after the applicable workload row’s displayed formula and GPU count fit your model and serving plan. Treat each row as a workload-specific estimate, then validate the stated configuration against your actual weights, KV cache, activations, sequence length, batch size, and runtime overhead before committing.

H100 SXM has published specifications of 3350GB/s memory bandwidth, 989 TFLOPS at BF16, and 1979 TFLOPS at FP8, while its 900GB/s NVLink figure describes interconnect rather than single-GPU speed. FP8’s 2x specified throughput relationship and lower bytes per parameter can change both fit and compute-bound speed, but these figures are specifications or modeled bounds, not measured application throughput; HBM generation changes capacity and bandwidth, not compute.

H100 SXM is not the right choice when the workload table requires a GPU count for which no eligible offer is shown or when a representative run fails the stated fit assumptions. In those cases, change the configuration or evaluate another GPU with verified capacity before comparing rental prices.

Turn the workload estimate into a rental decision

  1. Measure the memory peak

    Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.

  2. Match the sold configuration

    Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.

  3. Bound the experiment

    Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.

How Marlin helps

Marlin matches your workload requirements to the lowest-priced suitable GPU option across supported providers.

  • Broader coverage: Marlin matches across every supported provider, not only the providers priced on this page.
  • Unified comparison: Marlin compares supported providers in one matching process, reducing the need to check prices manually across separate sites.
  • Fit-aware choice: Marlin returns the cheapest option that satisfies your workload requirements, not simply the cheapest row regardless of fit.

Manually revisiting provider pages to compare rates and verify the $3.20 figure takes repeated effort; Marlin handles that comparison when matching your workload to the lowest-priced suitable option.

Before you rent

Storage: Check the linked storage or billing terms and the checkout quote to establish whether volumes and snapshots are charged separately from GPU time.
Checkpoint recovery: With spot capacity, work completed after the last checkpoint may need to be rerun after an interruption, increasing billed GPU time.
Launchable configuration: Verify the exact GPU count, topology and availability before scheduling the run. A memory estimate does not establish that the configuration can be rented.
Billing granularity: Per-second, per-minute, or per-hour rounding can make short runs cost more than the headline hourly rate implies; confirm the billing increment on the provider's own billing page before committing.

Workload memory and cost projections use the stated assumptions; validate them with measured memory and billed runtime. Prices were collected from provider pricing sources, while fixed hardware specifications came from the published GPU datasheet. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://vast.ai/pricing · https://www.hyperstack.cloud/gpu-pricing · https://datacrunch.io/pricing · https://www.runpod.io/pricing · https://crusoe.ai/cloud/pricing · https://www.together.ai/pricing · https://www.paperspace.com/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://lambda.ai/service/gpu-cloud · https://www.coreweave.com/pricing

Marlin beta

Stop comparing. Start running.

Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable option across supported providers.

Get started with Marlin