The Inference Capacity Wars: Who Actually Controls the AI Supply Chain
Hyperscaler capital expenditure guidance keeps climbing toward two hundred billion dollars a year at each of the largest cloud providers, open-weight models keep collapsing the cost of matching frontier quality, and enterprises keep discovering the same uncomfortable fact on rollout day: the constraint on their AI roadmap was never model quality. It is capacity. Tokens per second, at acceptable latency, in the jurisdiction the workload is legally allowed to run in, at a price that survives the next contract renewal. Model selection is a solved problem for most tasks. Getting enough of the model, reliably, is not.
The contrarian read on where this is heading: the AI supply chain is consolidating into a utility, and a utility has chokepoints. There are exactly three that matter, and none of them is the model. The first sits at the silicon layer: advanced packaging and high-bandwidth memory, not the GPU logic die itself, is the binding constraint on how many accelerators the world can actually build. The second sits at the datacenter layer: power and land, which do not scale on a chip company's roadmap, they scale on a grid interconnection queue. The third sits at the capacity layer: the hyperscaler and neocloud allocation desk that decides, when demand exceeds supply, whose workload gets served first. The strategic question for every buyer of AI is no longer which model to use. It is who controls your inference when demand spikes, and what happens to your roadmap when the answer is someone else.

The Layers Nobody Prices Correctly
The headline chip story, more GPUs, obscures the layers where the actual scarcity lives. Every accelerator ships as a package: the compute die, high-bandwidth memory stacked and bonded to it through advanced packaging, and the networking silicon that lets thousands of them act as one machine. HBM supply has trailed GPU logic supply for multiple product generations running, because the memory itself and the packaging processes that bond it to the compute die sit with a small number of specialized manufacturers whose capacity additions run on multi-year lead times, not a fab's quarterly output schedule. A GPU shortage headline is almost always, underneath, a memory or packaging shortage, and the fix is not "build more fabs," it is a slower, more constrained supply chain that a single chip company's roadmap cannot unilaterally accelerate.
The power wall is the layer executives most consistently underweight, and it is the one credible candidate for the binding constraint by 2027. A modern AI training or inference campus now requests grid interconnection at a scale that rivals a small city's total demand, and interconnection queues in the major buildout regions already run multiple years, a timeline that data-center construction, however well funded, cannot outrun. Land follows the same logic: sites need proximity to power, water for cooling, and fiber, and the inventory of sites that clear all three simultaneously is shrinking faster than the deal pipeline chasing it. Every dollar of capex guidance assumes the power gets delivered on schedule; the honest read of grid-interconnection data is that a meaningful share of the announced capacity is not going to arrive when the press release implied it would.
The third layer is where the buyer actually feels the first two: neoclouds versus hyperscalers versus sovereign capacity. Neoclouds, GPU-specialized providers that lease raw accelerator capacity without the full hyperscaler platform wrapped around it, exist because hyperscaler capacity is allocated first to the hyperscaler's own strategic customers and internal workloads, leaving a second market for buyers willing to trade platform breadth for availability and, often, price. Sovereign and regional capacity, the domestic or jurisdictionally-ring-fenced buildouts that have followed the AI capex wave, exist for reasons that are only partly economic: data residency law, national-security procurement rules, and a general unwillingness among governments to have their AI capability sit entirely on infrastructure another country controls. The three options are not substitutes for the same product; they are different trades on the same underlying scarcity, and the fact that a buyer can choose among them is itself the leverage the next section covers.

This is a Premium Article
Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.
Already a member? Sign in