RTX 4090 Rental Price: Cheapest Cloud per Hour (2026)
Choose RTX 4090 rental capacity for your workload using current offer terms and hardware-fit evidence drawn from 3 cited sources and updated 2026-10-07.
If the workload fits one RTX 4090 and requires uninterrupted capacity, choose Vast.ai at $0.36 per GPU-hour, with availability reported as In stock. Even if the workload can restart, no eligible spot offer is observed, so keep the on-demand path. If an eligible spot offer appears and the workload can restart, compare its current rate and measured job cost before switching.
Recommendation source: This on-demand recommendation is a marketplace quote, not a provider-published list rate. Recheck it before booking.
RTX 4090 is a GPU rented here as 1x RTX 4090. Across on-demand and spot terms, the GPU identity and hardware profile stay the same, so model state, activations, and runtime overhead must fit the configuration’s device memory.
Run a short, bounded representative job at the observed on-demand rate of $0.36 per GPU-hour after confirming the actual checkout rate, using the same model, precision, batch size, sequence length, and output-quality target planned for production; record peak memory, completed work, and billed runtime.
Current recommendation
For runs that cannot be interrupted, use Vast.ai at $0.36/hr (In stock).
The closest on-demand comparison is Vast.ai at $0.36/hr (In stock) versus RunPod at $0.74/hr (In stock): $0.38/hr, or $3328.80/year.
Today's prices
USD/hr · one observed offer per provider and pricing model
1 GPU
Every row below: Pricing: on-demand · Availability: In stock
What an hour buys
Computed from the recommended on-demand offer: Vast.ai at $0.36/hr for 1 GPU.
Hardware specifications
| Specification | RTX 4090 |
|---|---|
| Memory | 24 GB GDDR6X |
| Memory bandwidth | 1008 GB/s |
| Dense BF16 | 165 TFLOPS |
| Dense FP8 | 330 TFLOPS |
| NVLink | No NVLink |
| TDP | 450 W |
What your workload needs
VRAM needed is calculated from the formula shown in each row. A listed hourly cost appears only when a currently eligible on-demand offer exists for that exact GPU count. 3 listed workloads are omitted because no currently eligible on-demand configuration exists at the required GPU count.
Every row below: RTX 4090: 1 GPU · Listed hourly cost: $0.36/hr (In stock; observed 2026-10-07; price basis: Marketplace quote)
| Workload | VRAM needed |
|---|---|
| 7B Q4 inference | 4 GB 7B × 0.5 B (4-bit) × 1.2 (KV+activations) |
| 7B FP16 inference | 17 GB 7B × 2 B (FP16) × 1.2 (KV+activations) |
What the numbers say
RTX 4090 fit begins with the workload table’s displayed VRAM formulas and GPU counts, which estimate capacity using assumptions specific to each workload rather than one universal overhead. Use the matching row to screen the configuration, then run the intended model, batch size, sequence length, and serving settings to confirm peak memory and leave operational headroom before renting.
RTX 4090 lists 1008 GB/s of memory bandwidth for reasoning about bandwidth-bound decode and 165 BF16 TFLOPS plus 330 FP8 TFLOPS for reasoning about compute-bound training or prefill, but these are specifications rather than measured throughput. The FP8 peak is specified at 2x the BF16 peak, and FP8 can reduce bytes per parameter, yet actual fit and speed depend on model support and the software path. Test the same model, precision, batch size, sequence length, and output-quality target, then record completed tokens or training work, elapsed throughput, and billed runtime.
RTX 4090 does not fit when the workload’s measured peak memory exceeds its 24 GB of GDDR6X, when the required BF16 or FP8 path is unsupported by the software stack, or when a distributed run needs interconnect capabilities beyond its No NVLink status. Before choosing it, run the intended model and precision to record peak memory, and test representative multi-GPU communication when the workload spans devices.
Turn the workload estimate into a rental decision
- Measure the memory peak
Use the intended precision, batch size and sequence length. KV cache stores attention state during serving; activations and optimizer state depend on the training setup. A weight-only estimate leaves these out.
- Match the sold configuration
Read the GPU count next to the hourly cost. If no eligible offer exists for that count, the memory estimate is not a launchable rental. Check topology before splitting a job across GPUs.
- Bound the experiment
Set a test budget and stop condition. Record billed runtime and completed work at the confirmed checkout rate, including recovery time for a spot test.
How Marlin helps
Marlin matches your workload requirements to the lowest-priced suitable GPU option across supported cloud and GPU-cloud providers.
- Broader coverage: Marlin searches across every supported provider, including those without a listable offer on this page.
- One-pass comparison: Marlin compares supported providers within one matching process, reducing the need to check each provider’s pricing separately.
- Fit-qualified price: Marlin returns the lowest-priced GPU option that satisfies your workload requirements, rather than the cheapest row regardless of fit.
Manually maintaining this comparison means rechecking checkout prices and terms against the observed on-demand rate of $0.36 per GPU-hour as listings change; Marlin handles the provider comparison when matching a workload to the lowest-priced suitable GPU option.
Before you rent
Workload memory requirements, GPU counts, monthly and annualized costs, and specification-derived ratios are modeled rather than measured. Prices were collected from provider pricing sources, while fixed hardware figures came from published manufacturer specifications. Re-check current rates and availability on the linked provider pages and hardware figures in the manufacturer sources, then confirm the checkout terms. https://vast.ai/pricing · https://www.runpod.io/pricing · https://www.nvidia.com/en-us/geforce/graphics-cards/40-series/rtx-4090/
Marlin beta
Stop comparing. Start running.
Stop tracking prices manually and use Marlin to match your requirements to the lowest-priced suitable GPU option across supported providers.
Get started with Marlin