GPU Repatriation: The Utilization Math Deciding Where Inference Runs
For a decade, the question of whether to own infrastructure had a stable and boring answer. General-purpose cloud compute was in near-continuous deflation, the vendors passed enough of it through to keep renting attractive, and the class of workloads that genuinely belonged on owned hardware kept shrinking. Accelerators broke that pattern, and they broke it in a way that has not yet propagated into most infrastructure planning.
The reason is that accelerator pricing does not track the cost of production. It tracks scarcity. The binding constraints sit upstream in advanced packaging, high-bandwidth memory, power, and the allocation decisions made by a small number of suppliers and hyperscalers, a structure examined in the analysis of who actually controls the AI supply chain. When a rented price reflects scarcity rather than cost, the amortized cost of owned hardware becomes competitive far sooner than a decade of general-purpose compute economics would lead anyone to expect. That is the entire mechanism, and it is why this question is live for inference in a way it has not been live for application servers since roughly 2015.
The thesis is narrow on purpose. GPU repatriation is a utilization bet, the break-even utilization is computable in advance, and it is a materially different number from the one that governs general-purpose repatriation. Sustained inference at high utilization clears it comfortably. Bursty inference does not clear it at all, and gets worse rather than better with scale, because an idle accelerator is the most expensive idle asset in the modern stack. Nearly every organization arguing about this has not produced the number, and the argument cannot be settled without it.

Why Inference Is Structurally the Right Shape
The general repatriation case rests on a workload having flat, predictable demand, sufficient scale to amortize hardware, and infrastructure competence already on staff. Those conditions, the egress arithmetic, the people-cost line that repatriation stories habitually undercount, and the full workload taxonomy are the subject of the general treatment in the cloud repatriation question, and they apply here unchanged. What follows is the part that is specific to accelerators, where the arithmetic diverges.
A production inference workload serving steady traffic is close to an ideal repatriation profile on every general criterion. It runs continuously rather than in scheduled bursts. Its load is predictable within a band, because it is driven by product usage rather than by batch schedules. Its capacity requirement is knowable in advance from traffic data the organization already has. In economic terms it resembles a steady-state application tier far more than it resembles a batch analytics job, and steady-state application tiers have always been the canonical repatriation candidate.
What makes it newly interesting is that the premium being paid on the rented side is larger than the premium on general-purpose compute, and it is larger for a reason that will not resolve on a vendor's pricing roadmap. Scarcity rents persist as long as the constraint persists. An organization renting accelerators at scarcity pricing for a workload that runs flat at high utilization is paying an elasticity premium for elasticity it is not consuming, and paying it at a rate set by a shortage rather than by the cost of the underlying silicon.
There is a second driver that is not a cost argument at all and should not be smuggled into one. In a constrained market, the question stops being only what capacity costs and becomes whether it is available when demand rises. Owned capacity is guaranteed capacity. Organizations whose product roadmaps are gated on inference availability are increasingly buying certainty rather than savings, which is a legitimate justification with a completely different threshold. It should be argued on its own terms, in front of the people who own the roadmap, rather than folded into a spreadsheet where it will be mistaken for a cost saving that never materializes.

This is a Premium Article
Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.
Already a member? Sign in