H100 PCIe vs H100 SXM: Which GPU Should You Choose?

H100 PCIe and H100 SXM share a memory tier, but bandwidth, compute and the offered GPU topology can change which configuration works best.

· refreshed every 12 hours

Choose between H100 PCIe and H100 SXM using the rental topology and a representative benchmark. Their shared memory capacity does not settle runtime or price; use the current on-demand rates for the final cost comparison.

Run a bounded test with the same model, precision, batch size and sequence length on each feasible option. Record peak memory, billed runtime and completed work; use the actual checkout rate to compare compute cost per completed job.

Does the job depend on the bandwidth or inter-GPU links of the offered machine? Confirm the actual topology of each rental. If either configuration meets the memory and software requirements, benchmark both; the PCIe or SXM label alone does not establish the cheaper run.
Which option meets your deadline and job budget in a representative test? Measure completion time first. The H100 PCIe-to-H100 SXM modeled cost-per-job ratio is 0.82; it uses specification-based assumptions and does not establish real runtime or a deadline.
Can the exact configuration launch under the rental terms your job needs? Confirm the GPU count, topology and availability at checkout. Use spot only after testing checkpoint recovery; if the required configuration cannot launch, reassess fit before substituting another GPU.

Today's prices

USD/hr · 1x GPU · prices grouped by provider and purchase model

ProviderH100 PCIeH100 SXM
Vast.ai $2.00 $3.87
Hyperstack $2.50 · Not reported $3.20 · Not reported
RunPod $2.89 $3.49
DataCrunch Not offered $3.85
Crusoe Not offered $3.90 · Not reported
Together AI Not offered $3.99 · Not reported
Paperspace Not offered $5.95 · Not reported
Lambda Unavailable as of 2026-10-07 Unavailable as of 2026-10-07

Configuration prices used by workload rows

USD/hr · currently eligible on-demand configurations at the exact GPU count shown in the workload table

ProviderGPUConfigurationPricePrice basisAvailability
CoreWeave H100 SXM 8 GPUs $49.24/hr Published list rate Not reported
Lambda H100 SXM 2 GPUs $8.38/hr Published list rate In stock

see all H100 prices → · see all H100 prices →

Where the two parts differ

SpecificationH100 PCIeH100 SXM
Memory80 GB HBM2e80 GB HBM3
Memory bandwidth2000 GB/s3350 GB/s
Dense BF16756 TFLOPS989 TFLOPS
Dense FP81513 TFLOPS1979 TFLOPS
NVLink600 GB/s900 GB/s
TDP350 W700 W
Prices are refreshed every 12 hours. Marlin matches your requirements across supported providers. Try Marlin →
0.63×
H100 PCIe-to-H100 SXM on-demand hourly price ratio. Below one favors the first GPU on hourly rate; above one favors the second.
0.82×
H100 PCIe-to-H100 SXM modeled cost-per-job ratio. Below one favors the first under the stated model; validate with measured runtime.
$1460/mo
Continuous-use monthly baseline for one H100 PCIe at its current on-demand minimum; scale the budget to your billed hours.

What your workload needs

Find your parameter count on the model's Hugging Face card. Serving maps to the inference rows, training to the fine-tuning rows.

WorkloadVRAM neededH100 PCIeH100 SXMCheapest today
7B Q4 inference 4 GB 1-GPU configuration: $2.00/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H100 PCIe: 1 GPU at $2/hr
13B Q4 inference 8 GB 1-GPU configuration: $2.00/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H100 PCIe: 1 GPU at $2/hr
70B QLoRA 46 GB 1-GPU configuration: $2.00/hr (In stock; observed 2026-10-07; price basis: Marketplace quote) 1-GPU configuration: $3.20/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H100 PCIe: 1 GPU at $2/hr
70B FP8 training 154 GB No eligible 2-GPU configuration 2-GPU configuration: $8.38/hr (In stock; observed 2026-10-07; price basis: Published list rate) H100 SXM: 2 GPUs at $8.38/hr
70B FP16 inference 168 GB No eligible 3-GPU configuration 8-GPU configuration: $49.24/hr (Not reported; observed 2026-10-07; price basis: Published list rate) H100 SXM: 8 GPUs at $49.24/hr

What the numbers say

H100 PCIe and H100 SXM have the same recorded memory capacity, so a memory-fit question alone cannot distinguish them. Their bandwidth, compute specifications and interconnect differ. Confirm whether the offered machine exposes the inter-GPU links your distributed job actually uses.

H100 PCIe has an on-demand minimum of $2.00/hr and H100 SXM has an on-demand minimum of $3.20/hr. The H100 PCIe-to-H100 SXM on-demand price ratio is 0.63. When both configurations meet the same requirements, compare the current rates rather than assuming that either GPU is always cheaper. Spot listings have different interruption terms and are not the basis of that ratio.

Use a compute-bound training or prefill test and, for serving, a test at the intended sequence length and concurrency. A bandwidth or interconnect advantage can matter differently in each test. Choose the configuration that meets the measured deadline and total job budget; PCIe is not inherently the cheaper rental.

How to test the comparison

  1. Hold the workload constant

    Compare the same model, precision, batch size, sequence length and output-quality target. Changing these between GPUs changes the question being tested.

  2. Measure the result you need

    For serving, record latency and throughput at the intended concurrency. For training, record time for the same completed work. Check peak memory and failures in both cases.

  3. Price the complete run

    Multiply the full configuration's checkout rate by billed runtime. Compare storage, transfers and restart costs separately. A datasheet-based ratio is a screening model, not this measurement.

Before you rent

Data egress fees can raise the total bill; verify each provider's current outbound-transfer rate and included allowance before moving checkpoints or results.
Multi-GPU scaling efficiency can change total cost because synchronization overhead may add GPU-hours when a workload does not parallelize cleanly.
Interruptible instances can disappear mid-job, turning weak checkpointing into lost compute time and a longer completion window.
Billing granularity can make short experiments cost more than their runtime suggests; use each provider's pricing page to find the minimum charge applied to a brief rental.

Marlin matches your workload requirements to the lowest-priced suitable GPU across supported providers.

Memory fit and cost-per-job comparisons are models. Datasheet peaks do not establish application throughput; verify the listed configuration with your own workload. https://vast.ai/pricing · https://lambda.ai/service/gpu-cloud · https://www.coreweave.com/pricing · https://datacrunch.io/pricing · https://resources.nvidia.com/en-us-tensor-core/nvidia-tensor-core-gpu-datasheet · https://www.runpod.io/pricing · https://www.hyperstack.cloud/gpu-pricing · https://www.together.ai/pricing · https://crusoe.ai/cloud/pricing

Marlin beta

Stop comparing. Start running.

Marlin matches your workload to the lowest-priced suitable GPU across supported providers.

Get started with Marlin