Home » Blog » Cloud H100s Rent for $4 an Hour Now. Does Owning GPUs Still Pay? Cloud H100s Rent for $4 an Hour Now. Does Owning GPUs Still Pay?

Nvidia H100 GPU

Cloud H100s Rent for $4 an Hour Now. Does Owning GPUs Still Pay?

Two years ago, the GPU question answered itself. During the 2023–2024 shortage, H100 capacity often rented for $7 to $10 per GPU-hour when it could be found at all, and buying hardware was the only way to guarantee access. That world is gone.

On-demand H100 SXM instances now list at $3.99 per GPU-hour on specialized GPU clouds like Lambda, with other providers in the same range (RunPod around $2.69, Nebius near $3.15 on a one-year commitment). Marketplace spot capacity has occasionally dipped below $2, though spot is interruptible and both provider- and region-dependent, not a rate to plan a budget around. Hyperscalers cost more for the same silicon: AWS lists its 8×H100 p5.48xlarge at $55.04 per hour, which works out to about $6.88 per GPU-hour. Industry reporting estimated that AWS cut its effective H100 pricing by roughly 44% in June 2025, and the rest of the market followed.

Cheap rentals should have killed the case for owning GPUs. They didn’t. Used hardware prices fell just as hard, and the decision still turns on a single variable: how busy the GPUs stay.

The short version, up front: for steady, near-continuous work (sustained utilization above roughly 60%), buying a used server tends to win, and when the hardware runs close to flat-out it can pay for itself inside a year. For bursty or occasional work below about 30%, renting is cheaper and carries none of the operational overhead. The one factor most comparisons leave out, resale value, tilts the math further toward owning, because used data center GPUs still sell for real money while cloud spend recovers nothing. The rest of this piece shows the numbers behind those thresholds, at current market prices, and where each option fits.

What renting costs in mid-2026

Public price sheets make this the easy half of the equation. A current on-demand menu looks like this per GPU-hour: B200 SXM at $6.69, H100 SXM at $3.99, A100 80GB at $2.79. These vary by region and reservation model — reserved and cluster commitments run lower, while hyperscaler on-demand rates run meaningfully higher for the same silicon.

One thing that rate already includes is the rest of the computer. Lambda’s $3.99 H100 instance bundles the CPU, system RAM, and NVMe storage that surround the GPUs — its 8×H100 node ships with 1,800 GiB of RAM and 22 TiB of SSD, and AWS’s p5.48xlarge similarly pairs its eight H100s with 192 vCPUs, 2,048 GiB of RAM, and local NVMe. That matters for the comparison below: renting is a fully provisioned system, so a fair buy-side number has to be a complete server too, not eight bare accelerators.

An 8×H100 node at $3.99 per GPU-hour comes to about $31.92 per hour, or roughly $23,300 per month if it runs around the clock. That figure is the benchmark everything else gets measured against.

What owning costs

The used market is where ownership got interesting. New H100s have held at $25,000 to $40,000 each since mid-2024. Used AI servers have no transparent exchange, since configuration, OEM, networking, CPU generation, and warranty all move the price. Still, reported market transactions put a used, tested 8-GPU H100 server around $150,000 to $180,000, a fraction of what the same box cost new and far below the ~$500,000 sticker on a new Blackwell-generation server.

That range is for a complete server — GPUs, CPUs, RAM, storage, and networking already integrated — which is what makes it comparable to a rented instance. Buying eight bare H100 cards and building around them is a different exercise: the chassis, dual server CPUs, several hundred gigabytes of RAM, NVMe, and high-speed networking can add tens of thousands of dollars, and at 2026 component prices that gap has widened rather than shrunk. Whenever possible, price a fully built system against the cloud, not the accelerators alone.

Ownership carries operating costs that cloud pricing hides. A fully configured 8-GPU H100 server typically draws 8 to 10.2 kW, with NVIDIA’s DGX H100 specified at up to 10.2 kW at full load. The North American average for colocation was $196.25 per kW per month in the second half of 2025, per CBRE, though individual markets like Northern Virginia, Dallas, and Silicon Valley diverge from that average. At around 10 kW, that is roughly $2,000 a month for power and space alone. A realistic all-in figure is closer to $2,500 once connectivity, remote hands, and spares are added, though that extra ~$500 is a planning estimate rather than a quoted rate.

One warning from the current market: memory and storage prices rose sharply from late 2025 into 2026, so a used server that needs a DIMM or SSD refresh can cost more to bring online than the listing suggests. Careful buyers inspect first, or purchase from vendors that test and warranty their units.

The break-even math

The two options have different cost shapes. Owning is one upfront payment plus a roughly fixed monthly cost to run the server — a cost that barely moves whether the machine is flat-out or idle. Renting flips that: nothing upfront, but a bill that rises with every hour used. So utilization decides everything, because it sets how large the rental bill grows.

The method here is plain payback — purchase price divided by the amount owning saves each month, which gives the number of months to break even. Nothing fancier, and no depreciation or resale credit is baked in yet (that comes next, and it only helps the ownership case).

The example is one used 8×H100 server at $165,000, the middle of the reported range. Renting the same capacity costs about $23,300 a month flat-out. A server busy only half the time needs only half as much cloud, so the rental bill halves — but the owned server still costs its fixed ~$2,500 a month to keep running. Owning saves the difference between those two, and that shrinking difference is what stretches payback as utilization falls.

Utilization Cloud rental / month Owning saves per month (cloud bill − $2,500 opex) Payback on $165,000
100% (around the clock) ~$23,300 ~$20,800 ~8 months
50% ~$11,650 ~$9,150 ~18 months
30% ~$7,000 ~$4,500 ~37 months (3+ years)

Read down the table and the logic is plain: the busier the hardware, the faster owning it pays off. A server already bought costs about the same to run flat-out as it does sitting half-idle, while the cloud bill climbs with every hour. At 30% utilization the cloud simply costs less, with no hardware to operate.

This is one worked example, not a fixed answer. Change the purchase price, operating cost, or rental rate and the numbers move. Fill in the gap between the rows and the threshold becomes clear: at about 60% utilization the rental bill is roughly $14,000 a month, so owning saves about $11,500 and pays back in roughly 14 to 15 months, comfortably inside an 18-month planning horizon. That is the basis for the rule of thumb: sustained utilization above roughly 60% favors ownership, and below a third, renting wins. In between, the decision rides on the factors below and on an honest forecast, since teams routinely overestimate how busy their GPUs will be.

Buy the same capacity new and the picture shifts predictably. An 8×H100 HGX system runs about $250,000 to $320,000, with roughly $285,000 a typical mid-range quote, so the payback stretches by around 70% at every level: roughly 14 months at full utilization, past five years at 30%. The rental savings are identical; new hardware just costs about 70% more to capture them. That is why the example uses a used server. Unless a workload genuinely needs new hardware or a factory warranty, refurbished capacity captures the same economics for far less.

Resale value changes the equation

The break-even math above quietly assumes the hardware is worth nothing at the end. It isn’t, and this is the part of the calculation most published comparisons skip.

Data center GPUs have held economic value far longer than the “obsolete next generation” narrative suggests. Azure retired its V100 instances about 7.5 years after launch. One GPU cloud, CoreWeave, reported H100 capacity coming off 2022-era contracts being rebooked at 95% of original pricing — a single-operator data point, not market-wide evidence, but a striking one. Six-year-old A100s still trade actively at $7,800 to $18,900 depending on memory and condition.

Honesty requires the other side too. Used H100 pricing has been volatile (units that touched $50,000 at peak scarcity later sold at steep discounts), and the industry itself can’t agree on useful life. Amazon shortened its server depreciation schedule in 2025 while Meta lengthened its own in the same quarter, and in one widely covered short thesis Michael Burry argued the industry is understating depreciation by about $176 billion — his estimate and investment position, not an accepted industry figure. Nobody should plan around a precise resale number three years out.

But even a conservative assumption moves the needle. If that $165,000 server recovers a third of its cost at resale, the effective ownership cost drops by $55,000 and break-even arrives months sooner. Cloud spend recovers exactly nothing. This asymmetry is the strongest argument for ownership that cloud providers never mention, and capturing it just requires actually selling the hardware when upgrading instead of letting it depreciate in a storage closet. ITAD buyers such as BuySellRam.com purchase used GPUs and accelerators outright, which turns the residual value from a spreadsheet line into cash against the next generation of hardware.

Where cloud still wins

None of this makes ownership the default. Cloud remains the right call in specific, common situations.

Bursty or exploratory work is the obvious one. A team fine-tuning for two weeks a quarter should never own hardware for it. Access to the newest silicon is another: renting a B200 at $6.69 an hour beats a six-figure capital request when a project needs Blackwell-class memory bandwidth for a month.

Scaling and geography matter too. Serving inference in three regions, or absorbing a 10x traffic spike after a launch, is what hyperscale infrastructure is built for. And some teams simply shouldn’t take on hardware operations — a three-person startup has better uses for an engineer than debugging a failed power supply. GPU failure is not hypothetical: Meta reported 148 GPU failures across 16,384 H100s over 54 days of Llama 3 training. Annualized naively, that works out to roughly 6% of GPUs a year, and estimates run higher under sustained heavy load. Whatever the precise rate, hardware failure at cluster scale is a routine operational event, not a tail risk.

The hybrid default

Most teams past the experimentation stage end up in the same place: own the baseline, rent the burst.

Steady, predictable load, like the production inference that runs every hour of every day, goes on owned hardware, where high utilization makes the economics hard to beat. Spiky load (training runs, evaluations, traffic surges, anything regional) goes to the cloud, which charges nothing when idle.

The practical pattern for a growing AI company looks like this. Prototype on cloud until workloads stabilize. When a workload has run near-continuously for a quarter and shows no sign of stopping, price a used server against the trailing three months of cloud invoices; the invoice usually loses. Buy used or refurbished rather than new where the workload allows, since the discount is large and compute performance is effectively identical on healthy hardware. And when upgrading generations, sell the outgoing hardware promptly — residual values erode as each new generation ships, so timing the exit matters as much as timing the purchase.

A short decision checklist

Situation Better fit
Utilization under ~30%, or unpredictable Cloud
Sustained utilization above ~60% for 18+ months Own (used/refurb first)
Needs newest-generation hardware briefly Cloud
Steady production inference, cost-sensitive Own
Strict data locality or compliance Own or dedicated bare metal
Both steady base load and spiky peaks Hybrid

Prices in this article are July 2026 snapshots and move quickly — cloud rates have fallen for two years running, and used hardware reprices with every generation NVIDIA ships. The framework is what lasts: honest utilization measurement, residual value counted in the math, and GPUs treated as an asset with an exit price rather than a cost that vanishes.


When it is time to upgrade, the GPUs are not the only thing worth money. A retired AI server is a stack of resellable parts, and recovering all of it beats letting any of it sit in a closet. BuySellRam.com pays for the whole box: sell GPUs in bulk, sell memory, sell CPUs, and sell SSDs — the same components that made it a complete server in the first place. Request a quote and put that residual value back on the balance sheet.