RTX 4090 Rental Price: Cheapest Cloud per Hour (2026)

Choose RTX 4090 rental capacity for your workload using current offer terms and hardware-fit evidence drawn from 3 cited sources and updated 2026-10-07.

· refreshed every 12 hours

If the workload fits one RTX 4090 and requires uninterrupted capacity, choose Vast.ai at $0.36 per GPU-hour, with availability reported as In stock. Even if the workload can restart, no eligible spot offer is observed, so keep the on-demand path. If an eligible spot offer appears and the workload can restart, compare its current rate and measured job cost before switching.

Recommendation source: This on-demand recommendation is a marketplace quote, not a provider-published list rate. Recheck it before booking.

RTX 4090 is a GPU rented here as 1x RTX 4090. Across on-demand and spot terms, the GPU identity and hardware profile stay the same, so model state, activations, and runtime overhead must fit the configuration’s device memory.

Run a short, bounded representative job at the observed on-demand rate of $0.36 per GPU-hour after confirming the actual checkout rate, using the same model, precision, batch size, sequence length, and output-quality target planned for production; record peak memory, completed work, and billed runtime.

Does your model and serving workload fit the RTX 4090 capacity indicated by the adjacent table’s VRAM requirement and GPU count? If your workload matches FP8 or BF16 workloads within 24GB VRAM, validate the table’s workload-specific formula and displayed GPU count with a representative run that records peak memory. If the workload does not fit, select a configuration with sufficient aggregate device memory and suitable interconnect support. Before renting, require an observed offer for that exact configuration; the table’s modeled fit does not establish runtime or deadline performance.
Can your run checkpoint frequently enough to restart after an interruption? If your run can checkpoint and restart, no eligible spot offer is observed, so evaluate the on-demand option from Vast.ai at $0.36 per GPU-hour, with availability reported as In stock. If your run cannot tolerate interruption, keep that on-demand path. Before renting, confirm an observed offer for the exact configuration and measure whether it meets the runtime target.
At checkout, do the current on-demand price and availability facts change your choice between Vast.ai at $0.36 per GPU-hour with availability reported as In stock and RunPod at $0.74 per GPU-hour with availability reported as In stock? If an offer for the exact configuration remains available at checkout, use the supplied on-demand recommendation: Vast.ai at $0.36 per GPU-hour, with availability reported as In stock. In the observed comparison, Vast.ai at $0.36 per GPU-hour with availability reported as In stock is $0.38 per GPU-hour below RunPod at $0.74 per GPU-hour with availability reported as In stock, equivalent to a $3328.80 per GPU-year difference under continuous rental. If checkout rates or availability change, compare the current on-demand rates and measured job costs before renting.

Current recommendation

For runs that cannot be interrupted, use Vast.ai at $0.36/hr (In stock).

The closest on-demand comparison is Vast.ai at $0.36/hr (In stock) versus RunPod at $0.74/hr (In stock): $0.38/hr, or $3328.80/year.

Today's prices

USD/hr · one observed offer per provider and pricing model

1 GPU

Every row below: Pricing: on-demand · Availability: In stock

ProviderConfiguration $/hrPrice basis
Vast.ai $0.36 Marketplace quote
RunPod $0.74 Published list rate
2.06x
Use the observed same-GPU on-demand price spread to shortlist offers whose configuration and rental terms match the workload, then run the same representative job and compare completed work against billed runtime. A spot-to-on-demand cost comparison cannot be made from the observed evidence because no eligible spot rate is supplied.
unavailable
The spot-to-on-demand cost comparison cannot be made from the observed evidence because an eligible spot rate and a defined effective-cost ratio are not supplied. Verify an eligible quote for the exact configuration before comparing current rates, checkpoint overhead, restart time, and completed work.
$263/mo
The monthly baseline assumes full-time billing at a single tracked on-demand rate, while actual spending depends on billed GPU-hours. Replace that assumption with measured workload runtime and the actual checkout rate; compare on-demand with spot only when separate eligible quotes exist.

What an hour buys

Computed from the recommended on-demand offer: Vast.ai at $0.36/hr for 1 GPU.

$0.0150
Per GB of VRAM, hourly$0.36/GPU-hour divided by 24 GB of VRAM.
$0.0022
Per dense BF16 TFLOP, hourly$0.36/GPU-hour divided by 165 dense BF16 TFLOPS.
2800
GB/s of memory bandwidth per $/hour1008 GB/s divided by $0.36/GPU-hour.

Hardware specifications

SpecificationRTX 4090
Memory24 GB GDDR6X
Memory bandwidth1008 GB/s
Dense BF16165 TFLOPS
Dense FP8330 TFLOPS
NVLinkNo NVLink
TDP450 W

What your workload needs

VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 3 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.

Every row below: RTX 4090: 1 GPU · Listed hourly cost: $0.36/hr (In stock; observed 2026-10-07; price basis: Marketplace quote)

WorkloadVRAM needed
7B Q4 inference 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations)
7B FP16 inference 17 GB 7B × 2 B (FP16) × 1.2 (KV+activations)

What the numbers say

RTX 4090 fit begins with the workload table’s displayed VRAM formulas and GPU counts, which estimate capacity using assumptions specific to each workload rather than one universal overhead. Use the matching row to screen the configuration, then run the intended model, batch size, sequence length, and serving settings to confirm peak memory and leave operational headroom before renting.

RTX 4090 lists 1008 GB/s of memory bandwidth for reasoning about bandwidth-bound decode and 165 BF16 TFLOPS plus 330 FP8 TFLOPS for reasoning about compute-bound training or prefill, but these are specifications rather than measured throughput. The FP8 peak is specified at 2x the BF16 peak, and FP8 can reduce bytes per parameter, yet actual fit and speed depend on model support and the software path. Test the same model, precision, batch size, sequence length, and output-quality target, then record completed tokens or training work, elapsed throughput, and billed runtime.

RTX 4090 does not fit when the workload’s measured peak memory exceeds its 24 GB of GDDR6X, when the required BF16 or FP8 path is unsupported by the software stack, or when a distributed run needs interconnect capabilities beyond its No NVLink status. Before choosing it, run the intended model and precision to record peak memory, and test representative multi-GPU communication when the workload spans devices.

Turn the workload estimate into a rental decision

  1. Measure the memory peak

    Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.

  2. Match the sold configuration

    Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.

  3. Bound the experiment

    Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.

How Marlin helps

Marlin matches your workload requirements to the lowest-priced suitable GPU option across supported cloud and GPU-cloud providers.

  • Broader coverage: Marlin searches across every supported provider, including those without a listable offer on this page.
  • One-pass comparison: Marlin compares supported providers within one matching process, reducing the need to check each provider’s pricing separately.
  • Fit-qualified price: Marlin returns the lowest-priced GPU option that satisfies your workload requirements, rather than the cheapest row regardless of fit.

Manually maintaining this comparison means rechecking checkout prices and terms against the observed on-demand rate of $0.36 per GPU-hour as listings change; Marlin handles the provider comparison when matching a workload to the lowest-priced suitable GPU option.

Before you rent

Persistent storage: Verify any separately billed volume charge, billing unit, and retention period in the linked billing terms or checkout quote before committing.
Multi-GPU scaling: When the workload table requires several devices, total billed GPU-hours increase with the GPU count, while communication overhead can extend runtime; benchmark the exact configuration before projecting job cost.
Capacity confirmation: A collected availability observation is not a reservation, so verify that the exact GPU count and configuration can be provisioned at checkout before scheduling the run.
Billing granularity: Per-second, per-minute, or per-hour rounding can change the effective cost of short runs by billing unused portions of an interval, so confirm the increment and any minimum charge on the provider’s own billing page before committing.

Workload memory requirements, GPU counts, monthly and annualized costs, and specification-derived ratios are modeled rather than measured. Prices were collected from provider pricing sources, while fixed hardware figures came from published manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://vast.ai/pricing · https://www.runpod.io/pricing · https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/

Marlin beta

Stop comparing. Start running.

Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable GPU option across supported providers.

Get started with Marlin