Home » Blog » Tokens Just Got Cheaper Again. Will That Fix the AI Chip Shortage — or Fuel It? Tokens Just Got Cheaper Again. Will That Fix the AI Chip Shortage — or Fuel It?

On June 30, Anthropic released Claude Sonnet 5 at $2 per million input tokens and $10 per million output during its introductory window — less than half the $5 and $25 the flagship Opus 4.8 charges, while scoring 63.2 to the flagship’s 69.2 on the SWE-bench Pro coding benchmark. Roughly ninety percent of the capability at forty percent of the price.

That release is one data point on a curve that has been falling for three years. Stanford’s AI Index found that querying a model at GPT-3.5 level cost 280 times less in late 2024 than in late 2022 — from $20 per million tokens down to seven cents. Every major lab has been cutting prices on a similar slope, and open-weight models keep pulling the floor down further.

Note what is actually falling here: the price of tokens, not the price of AI. Total AI bills keep climbing — more on that below. Still, the token curve raises the question that matters to anyone buying hardware right now. If running a model keeps getting more efficient, shouldn’t that eventually take the pressure off GPU and memory demand? Shouldn’t DRAM, NAND, and accelerator prices come back down?

The short answer is no — cheaper tokens make the AI chip shortage worse, not better, and the rebalancing has to come from somewhere else entirely. The timeline for that is measured in years.

The intuition that keeps being wrong

The logic seems airtight. Cheaper inference means less compute per task. Less compute per task means fewer GPUs, less memory, smaller data centers. Efficiency should eat demand.

What actually happens is that cheaper inference changes which tasks are worth running at all. At 2022 token prices, putting a capable model behind every customer-service conversation was economically absurd. Layering sentiment analysis, escalation prediction, and quality scoring on top of each of those conversations was not even a budget line item. At 2026 prices, companies do all of it, routinely.

The spending data shows the result. Enterprise spending on generative AI hit $37 billion in 2025, up from $11.5 billion the year before — more than triple in a single year, and up from just $1.7 billion in 2023. Token prices fell by orders of magnitude over that same period. The total bill went up anyway, because usage grew faster than prices fell.

Token cost collapsed 280-fold while enterprise AI spending tripled.

A steam engine answered this question in 1865

The economist William Stanley Jevons watched exactly this pattern play out with coal. Watt’s improved steam engine burned far less coal per unit of work than its predecessors, and the confident prediction of the day was that England’s coal consumption would fall. Instead it soared. Cheaper steam power made steam economical in industries where it had never made sense before — textiles, railways, mining, shipping — and each new use created coal demand that had not existed.

Jevons published the observation in The Coal Question, and 160 years later it describes the AI hardware market almost line for line. Every time inference gets cheaper, some previously unprofitable application crosses the threshold into viability, and the new applications consume more total compute than the efficiency gains saved.

How cheaper tokens end up raising memory prices

What that means, chip by chip

The demand created at the software layer lands, physically, on a small number of component markets. They are not absorbing it gracefully.

DRAM. Server DRAM contract prices — the rates large buyers negotiate directly with manufacturers, which reach retail shelves with a lag — rose a record 90 to 95 percent in the first quarter of 2026, and TrendForce’s follow-up forecast calls for another 58 to 63 percent in the second quarter. The mechanism is simple: memory makers are shifting wafer capacity toward servers and toward HBM (high-bandwidth memory, the stacked DRAM mounted directly beside AI processors), while cloud providers lock up supply through long-term agreements.

NAND and enterprise SSDs. Same story, slightly behind on the curve: NAND contract prices are projected to climb 70 to 75 percent in Q2, with output increasingly steered toward enterprise SSDs for data centers while consumer product lines get thinner allocations.

GPUs. Here the pressure shows up less as sticker price and more as time. SK Hynix has sold out its entire 2026 output — HBM, DRAM, and NAND alike, and Micron has signaled that its 2026 HBM allocation will sell out as well. The other choke point is CoWoS, TSMC’s advanced packaging process — the step that physically bonds a GPU die to its HBM stacks. Packaging runs on its own production lines with its own capacity ceiling, which is why chips that have already been fabricated can still sit waiting: there is nowhere to assemble them. That is the scale Goldman Sachs’ model reflects when it projects $7.6 trillion in cumulative AI infrastructure investment — compute, data centers, and power combined — through 2031.

CPUs. The partial exception, and worth naming. A typical AI training server pairs eight GPUs with two CPUs; the CPU is simply not the scarce part, and CPU pricing has stayed comparatively calm. But general-purpose server buyers are not off the hook — memory now makes up a growing share of any server’s bill of materials, so the total cost of even a CPU-centric box is rising with DRAM.

Anyone who lived through the 2017–2018 crypto-mining GPU shortage or the pandemic-era chip crunch might expect this episode to end the same way: demand evaporates or logistics normalize, and prices snap back. The structure this time is different. Those were demand shocks hitting a fixed supply. This one is a deliberate reallocation of manufacturing capacity, locked in by multi-year supply agreements between memory makers and cloud providers — the buyers cannot easily walk away, and the capacity cannot easily walk back.

Why software efficiency won’t rescue hardware prices

Optimization keeps arriving: speculative decoding (a small model drafts text that the large model merely verifies), quantization (running models at lower numerical precision), better GPU kernels (hand-tuned low-level code). Each one cuts the compute needed per token, and each one gets briefly read as the beginning of the end of the shortage.

Meta’s response to the memory crunch is the clearest counterexample. Rather than discard the DDR4 memory inside its decommissioned servers, Meta built a custom controller that lets those retired modules keep serving live AI workloads alongside brand-new machines — cutting the server count needed for inference by up to 25 percent. Note what that is: one of the world’s largest infrastructure operators going to considerable engineering lengths to reuse years-old memory, because new DRAM is too expensive and too scarce to buy. Efficiency did not shrink Meta’s appetite. It sent Meta digging through its own retirement pile.

Software optimization compounds monthly. Fab construction is measured in years, and no amount of clever code shortens it. The efficiency gains flow into more usage; the physical constraints stay physical.

The counterargument — and why it falls short

There is a serious version of the opposing case. Custom AI chips (ASICs) run specific workloads more efficiently than general-purpose GPUs. Mixture-of-experts architectures activate only a fraction of a model per query. Small models keep closing the gap with large ones. Stack those trends and total hardware demand could, in principle, flatten.

Two problems. First, this efficiency revolution already happened — that is what the 280-fold cost decline was — and hardware demand tripled during it. Model-level efficiency is the fuel of the paradox, not the cure. Second, Goldman’s analysts ran the ASIC-substitution scenario directly: if compute demand were fixed, cheaper chips would shrink total investment; if demand is elastic — buyers consume more as it gets cheaper — substitution mainly redistributes revenue among chip suppliers without shrinking the total. Three years of evidence point firmly to elastic. And ASICs offer no escape from the tightest bottleneck anyway: they need the same HBM and the same advanced packaging as the GPUs they would replace.

So what actually rebalances supply and demand?

If efficiency feeds demand instead of relieving it, only two mechanisms can close the gap.

The first is new supply, and its progress is measurable. TSMC is expanding CoWoS packaging capacity aggressively, yet industry reporting as of June still puts demand roughly 20 percent above supply, narrowing to perhaps 10 percent by the end of 2026 — every expansion so far has been absorbed before it landed. On the memory side, TrendForce expects no meaningful DRAM capacity growth before late 2027. New fabs are the only force that genuinely lowers prices, and the fabs that would do it are still under construction.

The second is demand getting priced out, and there are early signs of it. The latest outlook shows memory price increases decelerating as buyers hit an affordability ceiling — DRAM still rising 13 to 18 percent in Q3 2026, but no longer doubling by the quarter. That is rebalancing of a sort, just not the comfortable sort: it means PC and phone buyers stepping back so data centers can keep buying.

The realistic picture: memory stays expensive into late 2027 and plausibly 2028, GPU availability is governed by packaging queues into 2027, and CPUs remain the one calm corner of the market. Cheaper tokens accelerate none of this relief. They delay it.

The consumer side of the squeeze

None of this stays inside the data center. TrendForce’s Q1 forecast put PC DRAM contract prices up 105 to 110 percent in a single quarter — the steepest increase on record — and reports that smartphone and notebook brands have begun raising prices and quietly downgrading memory specs to cope. Dell raised system prices 15 to 20 percent in mid-December, with Lenovo following in January, and smaller builders feel it hardest: Framework has been posting monthly updates warning customers about memory and SSD costs through 2026.

For anyone weighing a PC build, a laptop refresh, or a phone upgrade, the practical translation is blunt: memory and storage costs are baked into consumer hardware for at least another year, and waiting for a return to 2025 prices means waiting for those 2027 fabs, not for the next sale.

The other side of the ledger: retired hardware is appreciating

For most organizations, this market creates two very different problems. Buyers face acquisition costs that have doubled in quarters, not years. But the same shortage cuts the other way: businesses retiring older infrastructure are discovering that hardware written off as obsolete now holds real resale value.

Meta’s DDR4-recycling project is the loudest possible signal that older memory has genuine second-life demand — server DIMMs outlast the servers they ship in, and the market has noticed. Decommissioned server memory, workstation RAM, and enterprise SSDs are trading at multi-year highs while new-part prices climb.

For organizations with retired equipment on shelves, this is an unusually good window to sell used server memory or offload enterprise SSDs and drives rather than letting them sit through the one period in years when they have been appreciating instead of depreciating.

The window has a shape. Until new fab capacity arrives in late 2027 at the earliest, every token-price cut from the model labs feeds more usage, more usage feeds more hardware demand, and the value of working memory — new or used — keeps being set by a market that cannot build fabs any faster.