When Models Commoditize, Where Does the Margin Go?
Every generation of software has had a layer that looked like the business and turned out to be the commodity. Databases, operating systems, and cloud compute each went through a phase in which the product at the centre of attention commanded premium prices, followed by a phase in which the value moved to whoever sat above or below it. The large language model is going through the same transition, only faster than any layer before it.
The thesis of this piece: the model layer is commoditizing on a timescale of months rather than years, and the margin is moving to the two ends of the stack. At one end sit the owners of scarce inputs: compute, energy, proprietary data, and distribution. At the other sit the owners of the workflow and the customer relationship. The middle, meaning thin wrappers and undifferentiated model resale, is being squeezed from both sides.
This is not the same as saying models are worthless. The frontier still commands a premium, and it will for as long as a meaningful gap exists at the hardest tasks. But the premium attaches only to the frontier, the frontier moves every few months, and everything behind it falls in price toward the cost of serving it. For an executive, the practical questions are where the durable margin sits, what that means for how AI is bought, and where a software company should build. Answering them starts with the evidence.

A Fixed Capability Gets Cheap Fast; the Frontier Does Not
The single most important fact about model pricing is that two different things are happening at once, and most commentary blurs them.
The price of any fixed level of capability is collapsing. Epoch AI measured how quickly the price of reaching a given benchmark score fell across six benchmarks over three years, and found declines ranging from 9x to 900x per year depending on the task, with the price of GPT-4-level performance on PhD-level science questions falling about 40x per year (Epoch AI, March 2025). Stanford's 2025 AI Index found that the inference cost of a system performing at the level of GPT-3.5 dropped more than 280-fold between November 2022 and October 2024 (Stanford HAI). Andreessen Horowitz, comparing the cheapest model at each capability level over three years, put the rate at roughly 10x per year and called it "LLMflation" (a16z).
Open-weight models keep closing on the frontier. Epoch AI reported in May 2026 that since January 2026 the most capable open-weight models have trailed the best closed models by an average of about four months on its capabilities index (Epoch AI, May 2026). Four months is less than a typical enterprise procurement cycle. Whatever a closed model can do today, a model that anyone can download and serve will do before most buyers have finished evaluating the first one. What open weights do to the licensing economics specifically is examined in the open-weights endgame.
The price of the frontier itself has barely moved. GPT-4 launched in March 2023 at 30 dollars per million input tokens and 60 per million output tokens (OpenAI developer forum, March 2023). OpenAI's current flagship lists at 10 and 50, while its cheapest current text model lists at 10 cents and 50 cents, a hundredfold spread within a single catalogue (OpenAI API pricing, September 2026). Anthropic's catalogue shows the same shape: its top model lists at 10 and 50 per million tokens, its smallest current model at 1 and 5. Within one product tier, Anthropic's Opus line has fallen from 15 and 75 for Opus 4.1 to 4 and 20 for Opus 5.5, and the introductory price on Sonnet 5 was made permanent rather than raised as originally scheduled (Anthropic pricing, September 2026).
Put those together and the shape of the market is clear. The frontier is a premium product with a short shelf life. Each frontier model is a premium product for a few months, then a mid-tier product, then a commodity, and the whole cycle takes a year or two. A business built on reselling model capability is selling an asset that depreciates by an order of magnitude per year. A business built on something the falling price does not touch can treat that depreciation as a subsidy. The rest of this piece is about which is which.

This is a Premium Article
Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.
Already a member? Sign in