Cloud
GPU Cloud Pricing Explained: On-Demand, Spot and Reserved
How GPU cloud pricing models work, why the cheapest hourly number can be misleading, and how to compare on-demand, spot and reserved GPU capacity.
Updated
GPU cloud pricing looks simple when every provider shows a dollar-per-hour number. In practice, two prices that look similar can represent very different products.
The most important distinction is the purchasing model: on-demand, spot or interruptible, and reserved or committed capacity.
On-demand GPU pricing
On-demand capacity is the flexible option. You start a GPU when you need it and stop paying when the provider's billing rules say the instance has ended.
It is useful for experiments, short training runs, temporary inference capacity and teams that do not yet know their long-term utilisation.
The trade-off is price and availability. On-demand capacity is usually more expensive than a long-term commitment, and the newest accelerators can be unavailable in a preferred region even when a provider advertises them.
Spot and interruptible GPUs
Spot capacity uses spare infrastructure that the provider can reclaim.
The lower hourly rate can be valuable for fault-tolerant workloads, batch processing, rendering, hyperparameter searches and training systems that checkpoint frequently.
The risk is interruption. A cheap GPU that disappears halfway through an uncheckpointed job can become more expensive than a stable instance once lost compute time is included.
Before using spot capacity, check how termination warnings work, whether pricing changes during the instance lifetime, and how quickly the workload can recover.
Reserved and committed capacity
Reserved capacity trades flexibility for predictability.
A provider may offer lower rates when a customer commits to a month, several months, a year or longer. Large deployments can also be negotiated privately rather than purchased from a public rate card.
This model fits teams with steady utilisation and production workloads where guaranteed access matters more than instant flexibility.
The important point is that a public on-demand price and a negotiated reserved price are not directly comparable.
Compare cost per useful unit of work
Hourly price is only the first layer.
A faster GPU can cost more per hour but complete the workload sooner. A system with better networking can scale across eight GPUs more efficiently. More VRAM can let a model run on fewer devices. A provider with lower egress charges may be cheaper over the full project even if its GPU rate is higher.
Useful comparison metrics include:
- total cost of a training run
- cost per million tokens served
- cost per completed batch
- time to train
- effective utilisation
- storage and data-transfer cost
- cost of idle capacity
- cost of failed or interrupted work
The GPU cloud cost calculator can help turn an hourly rate into a project-level estimate.
Make sure the hardware is comparable
GPU names alone are not enough.
An H100 listing may refer to PCIe, SXM or NVL hardware. Multi-GPU servers can have different interconnects. Regions may use different instance designs. Some marketplace listings include consumer-style host configurations while others are dedicated data-center systems.
For H100 specifically, see H100 PCIe vs SXM. For generation choices, compare H100 vs H200 and B200 vs H200.
Freshness matters
GPU prices change as supply expands, new accelerators launch and providers adjust capacity.
A comparison without a source date can quickly become misleading. GPU Data Hub's GPU Price Tracker shows a last-updated timestamp for every row and leaves prices unverified when a provider-published figure cannot be confirmed.
Provider rate cards should still be checked before a purchase because availability, taxes, regions and contract terms can change.
A practical buying process
Start with the workload: model size, required VRAM, expected run time and whether multiple GPUs need high-bandwidth communication.
Then choose the accelerator class. Compare providers that actually offer that accelerator in a suitable region. Separate on-demand, spot and committed rates instead of mixing them into one ranking.
Finally, calculate total workload cost rather than simply picking the lowest hourly number.
Browse the current GPU Price Tracker, compare GPU cloud providers, or explore the GPU database for hardware specifications.
GPU Data Hub Daily
Stay Ahead of the AI Infrastructure Economy
The most important GPU, AI, data-center, semiconductor and cloud developments delivered directly to your inbox.
Cite this page
“GPU Cloud Pricing Explained: On-Demand, Spot and Reserved.” GPU Data Hub. https://gpudatahub.com/guides/gpu-cloud-pricing-on-demand-spot-reserved
You are welcome to reference and link to this page. Please link to the URL above.