On-demand vs reserved: choosing a GPU pricing model
Reserved capacity is meaningfully cheaper per hour, and on-demand is 100% cheaper when idle. The only question that matters is utilization — here's the break-even math.
Maya Reyes
CEO, Reviosa
May 28, 2026 · 6 min read
Two prices, one question
Every GPU cloud offers some version of the same trade: pay more per hour with zero commitment, or commit to a term and pay less per hour whether you use it or not. On Reviosa, on-demand H100s are $2.49/hr with per-second billing; reserved commitments price 25-40% below on-demand depending on term length and volume. Vendors love to make this choice sound complicated. It isn't. It is a single-variable problem, and the variable is utilization.
The break-even math
Take a single H100 with a reservation priced 30% below on-demand — about $1.74/hr effective, or $1,272/month. On-demand at $2.49/hr costs exactly what you use. The crossover is mechanical:
| Utilization | On-demand monthly | Reserved monthly (30% off) | Winner |
|---|---|---|---|
| 25% | $454 | $1,272 | On-demand |
| 50% | $909 | $1,272 | On-demand |
| 70% | $1,272 | $1,272 | Tie |
| 100% | $1,818 | $1,272 | Reserved |
The rule generalizes: break-even utilization equals one minus the reserved discount. A 30% discount breaks even at 70% utilization; a 40% discount at 60%. If you can honestly forecast a GPU busy more than that, reserve it. If you can't, don't.
Note what is not in the formula: team size, model size, how fast you're growing, or how the last negotiation went. Utilization is the entire equation, which is why measuring it honestly matters more than any pricing spreadsheet.
Why teams get this wrong
The most common failure mode is reserving based on peak demand rather than average demand. A team that trains hard two weeks per month is at ~50% utilization — solidly on-demand territory — but the memory of waiting for capacity during those two weeks drives them to reserve, and the idle half of the month quietly erases the discount. The second failure mode is the inverse: running a 24/7 inference fleet on on-demand pricing out of inertia, donating 30% month after month for flexibility they never exercise.
Per-second billing sharpens the math in on-demand's favor for bursty work. When a 9-minute CI job on an 8×H100 node costs $2.75 instead of a full billed hour, "spiky" workloads get dramatically cheaper — and the utilization bar a reservation must clear gets higher.
The hybrid pattern that usually wins
Almost every mature customer converges on the same structure:
- Reserve the floor. The inference fleet's overnight minimum, the always-on dev node, the recurring nightly training window — capacity you can predict a quarter out.
- Burst on demand. Experiments, hyperparameter sweeps, evals, load tests, and the traffic above your inference floor.
- Review quarterly. Pull your utilization report, compare your reserved floor to your actual trough, and adjust at renewal. Fifteen minutes, four times a year.
A concrete example: Loomline AI runs six reserved H100s as their serving floor and bursts to 10-14 on-demand GPUs during business hours. Their blended rate lands ~22% below pure on-demand with zero capacity anxiety — better than either pure strategy would deliver at their traffic shape.
How to decide in five minutes
- Pull last quarter's GPU-hours from your billing dashboard (Reviosa itemizes this per instance).
- Divide by the hours in the quarter to get true utilization per steady workload.
- Anything above ~70%: get a reservation quote. Anything below: stay on-demand.
- Re-run the numbers when workloads change — a product launch or a new training cadence moves the answer.
The honest version of this advice sometimes costs us revenue: plenty of teams should not reserve, and we will tell you so. Idle committed capacity helps nobody — we would rather you buy exactly what your utilization justifies and stay for years than overcommit and churn.
Run it yourself
Same fleet, your workload
Everything we write about runs on hardware you can rent by the second — H100 SXM from $2.49/hr, live in under 90 seconds.
$10 free credit for new accounts · Per-second billing · No egress fees