Pantheon
Back to Research

The H100 Cliff That Isn't: What the GPU Resale Market Is Actually Telling Buyers

Milo
8 min read
Rows of NVIDIA data-center GPUs in a server rack, representing the AI compute secondary market

If you've read a tech-finance headline in the last six months, you've been told a GPU glut is coming. Used H100s collapsing in price. Hyperscalers cooking their books on depreciation. The whole AI capex cycle one quarter away from a reckoning. It makes for a great short thesis and a worse buying guide.

We sit on the other side of those headlines. We source both ends of this market — new GB300 racks coming off the OEM line and used H100s coming off expired contracts — and we watch the deal flow that the index charts only summarize. From here, the "glut" story is mostly a category error. What people are calling oversupply is a generational rotation to Blackwell, and that is a very different thing from demand falling off a cliff. The distinction is the whole ballgame for anyone actually writing a check this year.

What the rental chart actually shows

Start with the number everyone points at. H100 cloud rental did collapse. In 2023 you paid something like $8 an hour for an on-demand H100, with spot peaks well past $12 when the scarcity was at its worst. By late 2025 the marketplace median had fallen toward ~$2/hr (Silicon Data's H100 rental index), and AWS cut its on-demand H100 pricing by roughly 30% mid-year. If you stop reading there, "glut" is the obvious conclusion.

But the chart didn't keep falling. It bounced. Rental rates ticked back up into 2026 — the index climbed off its lows toward the $2.20–$2.35/hr range, and one-year contract pricing rose almost 40% off an October 2025 floor (Silicon Data). A market in genuine oversupply does not firm up four months after bottoming. What you're looking at is a price normalizing after a two-year scarcity premium burned off — the H100 settling into what it's actually worth as an inference workhorse, not a collapse toward zero.

The used-card market tells the same story in a different key. Used and refurbished H100s changed hands as high as ~$50K during the mid-2024 crunch; by late 2025 they'd fallen to roughly the mid-$20Ks to mid-$30Ks depending on condition and whether the unit was refurbished or simply used (Compute Exchange, IntuitionLabs). That's a big drop. It is also exactly what you'd expect when the next generation starts shipping in volume and the people who waited for it stop bidding up the old one. The H100 didn't get worse at its job. A better card showed up, and the scarcity tax expired.

Watch the spread, too. Refurbished units have been clearing in the mid-80% range against new pricing while genuinely used cards trade at a materially deeper discount — which tells you the market is pricing condition and provenance carefully, not dumping a commodity. That's the behavior of a functioning secondary market finding fair value, not a fire sale.

Why this is rotation, not collapse

Here's the part the glut narrative quietly omits: demand is not shrinking. It's moving.

Blackwell is projected to make up over 70% of NVIDIA's high-end GPU shipments in 2026, up from around 61% the year before (TrendForce). The constraint upstream isn't soft orders — it's that HBM memory production and TSMC's CoWoS packaging capacity are sold out well into 2027. Every serious buyer who deferred an H100 purchase in 2024, reasonably betting on the next architecture, converged on the Blackwell line at the same moment the hyperscalers did. That's not a market losing interest in compute. That's a market stampeding one generation forward and leaving the previous one to reprice.

So the H100 going cheaper isn't a demand signal. It's a vintage signal. The frontier moved, and the trailing edge got more affordable — which, if you're buying inference capacity rather than training the next frontier model, is the best news you've had in two years.

The depreciation panic, sized correctly

The other half of the bear case is an accounting argument, and it deserves a real answer rather than a dismissal. In November 2025, Michael Burry publicly accused the hyperscalers of understating GPU depreciation — booking these chips over five-to-six-year useful lives when the real economic life is closer to two or three, a gap he pegged at roughly $176 billion of understated depreciation across 2026–2028 (CNBC). For specific balance sheets carrying enormous fleets at aggressive book lives, that's a genuine concern, and we won't wave it away.

But "the accounting is generous" and "the hardware is worthless in three years" are two different claims, and the second one isn't holding up. CoreWeave's CEO has said that when a 2022-vintage GPU contract expired, the hardware was immediately re-leased at around 95% of its original rate (CNBC). Three-year-old silicon clearing at 95% of its first-day price is the opposite of a cliff. It's durable value — older accelerators finding a productive second life on inference workloads that don't need the bleeding edge.

Both things are true at once, and a buyer should hold them together: the depreciation debate is real for certain over-levered operators, and the depreciation cliff — the fleet-wide melting-ice-cube story — is overstated. A used H100 retains real utility for years. What it doesn't retain is a scarcity premium that was never structural to begin with.

The signal from where there's no secondary market at all

One more tell, and it's the one that only shows up if you're actually in the supply chain. There is essentially no secondary market for GB300 or Blackwell-Ultra parts — none, because not a single one has come off contract. The GB300 NVL72 — the current frontier-training rack, 72 Blackwell GPUs lashed together with 36 Grace CPUs — moves entirely through primary cloud and OEM channels, and it's spoken for well into 2026. You can't buy a used one because nobody is done with theirs.

That's your demand signal, cleaner than any rental index. If the AI buildout were genuinely oversupplied, the newest, most expensive hardware would be the first thing to soften. Instead it's the hardest thing in the world to get. The cheap end of the market is cheap because it's last year's, not because nobody wants compute. Even the gray market agrees: with U.S. export controls tightening, an estimated $1B+ of NVIDIA chips were smuggled into China in 2025, and prices for five-year-old A100 servers there have tripled (Tom's Hardware, CNBC). Hardware nobody wants doesn't get smuggled.

What an operator should actually do

Strip out the headlines and the buyer's read is unusually clean this year:

  • Buying for frontier training? Buy new, and buy Blackwell — GB300-class racks through primary channels. There's no used shortcut to the frontier, and the secondary market can't help you here because it doesn't exist yet.
  • Buying for inference, fine-tuning, or steady production load? Used and refurbished H100 pencils out if you don't overpay. The card is cheaper for structural reasons, not because it stopped working. Anchor to the marketplace median, discount for vintage and condition, and don't let a panicked seller's framing — or a hopeful one's — set your price.
  • Watching the depreciation debate? Treat it as a balance-sheet question about specific operators, not a verdict on the asset class. The hardware holds value. Some books are carried optimistically. Both can be true.

The mistake we see most often is treating the rental chart as a referendum on whether to buy compute at all. It isn't. It's a repricing of which compute, and the answer depends entirely on what you're running. The glut headline is selling you a market collapse. The deal flow is selling you a better entry point on inference and a sold-out frontier — which is a much more useful thing to know.

This is the compute half of a two-part story. The other half is power: even a perfectly-priced fleet of GPUs is dead weight without megawatts behind the meter to run it, and that market has its own contrarian read — the power side of the same buildout. Power and compute are the two halves of every real project that crosses our desk, and they're almost never solved by the same person. The third piece in this series closes the loop on why: the binding constraint isn't chips or megawatts but time-to-energized — and you don't own your compute roadmap unless you own your power timeline.

We source both new GB300 racks and secondary H100 supply, and we watch this deal flow closely — but we don't set the market, and we publish only public price data, never our own quotes. If you're sizing a buy this year and want a straight read on new-versus-used for your actual workload, browse the catalog or talk to us.


Pantheon Research is our series on the infrastructure behind AI: power, turbines, cooling, compute, and the procurement reality that decides who actually ships. Field notes from the deal flow, not the keynote.

Tags:
research
gpu
compute
ai-infrastructure

Milo

Expert in manufacturing technology and industrial solutions, sharing insights on the latest trends and best practices.

Share this article

Related Articles

Continue exploring manufacturing insights and industry expertise

Rows of NVIDIA data-center GPUs in a server rack, representing the 2026 AI accelerator buying decision
7/8/2026

Which NVIDIA GPU Should You Buy in 2026? H100 to GB300 NVL72

A field note from the deal flow: how to pick the right NVIDIA accelerator in 2026, from Hopper H100 and H200 up through Blackwell B200, Blackwell Ultra B300, and the GB300 NVL72 rack. No spec-sheet worship and no price theater — just how we actually help operators choose by workload, memory footprint, availability, and what their facility can cool.

Read Article
Industrial high-desert testing ground in Northern Nevada, representing physical proving grounds for embodied AI and robotics
7/9/2026

Robots Run on the Same Stack: Embodied AI Is a Power-and-Compute Story

The humanoid headlines are all about the machine — the hands, the gait, the demo. From where we sit, the machine is the least interesting part. A robot fleet is a distributed data center that happens to move: it learns on the same GPUs, runs on the same electrons, and needs a physical place to fail safely. Which is why embodied AI is quietly an infrastructure decision — and why it keeps pointing at the same desert as everything else.

Read Article
A data-center and industrial power operations floor, representing the software layer that runs and reads physical assets
7/10/2026

Industrial AI Software Runs Next to the Iron: The One AI Layer You Can't Build Remotely

Most AI software is written wherever the engineers happen to live. Industrial AI — the grid optimizers, the predictive-maintenance models, the digital twins — is the exception, and operators keep learning it the hard way. It's only as good as its proximity to real physical assets, the telemetry those assets exhaust, and a real industrial customer next door. From the desk where we source the iron, this is the software layer we watch bolt itself to the hardware — and it's the reason the fourth vertical in this series has the same address as the first three.

Read Article

Ready to Transform Your Manufacturing?

Discover how our advanced machinery can improve your production efficiency and product quality.