Frontier Model Lock-In: The Hidden Switching Costs of Betting on One AI Provider

Frontier Model Lock-In: The Hidden Switching Costs of Betting on One AI Provider

The procurement team that standardized on a single frontier model provider eighteen months ago did the rational thing. One provider meant one contract, one set of credentials, one billing relationship, one SDK to learn, and one mental model for the engineering organization to internalize. Velocity is real, and in the early phase of an AI program velocity is the entire game. The team shipped. The reflex that produced the decision was sound, and the decision looked, on the procurement comparison that justified it, almost free to reverse. The pitch was that the model is behind an API, the API speaks a near-standard format, and switching providers later would be a base-URL change and a key swap. A weekend of work, on the slide.

The slide was wrong, and the organizations now trying to leave are discovering by how much. The switching cost of a frontier model provider is not the integration. It is everything that accreted on top of the integration once the model became load-bearing, and almost none of it was visible at procurement time because almost none of it existed yet. The contrarian thesis is straightforward and increasingly hard to argue with: frontier model lock-in is the new cloud lock-in, and like cloud lock-in it hides not in the obvious contract terms but in the thousand downstream dependencies that form silently as a platform becomes essential. The difference is that the industry learned to see cloud lock-in over a decade of painful migrations, and it has not yet learned to see model lock-in, which means most organizations are accumulating it right now without measuring it.

Single vs multi

Where the Lock-In Actually Lives

The base-URL-swap mental model assumes the model is a fungible component behind a stable interface. The reality is that a frontier model is not a component. It is a behavior, and the organization adapts itself to that specific behavior in layers that compound. The evidence from teams that have attempted a provider migration shows the cost concentrating in five places, none of which appear on a price comparison.

The first is prompt and tool-call formatting tuned to one model's quirks. A mature AI feature is not one prompt. It is dozens to hundreds of prompts, each iterated against the behavior of a specific model, each carrying small accommodations for that model's tendencies: how it handles system instructions, how literally it follows formatting demands, how it responds to few-shot examples, how its tool-calling schema wants arguments shaped. None of this is portable. A prompt that produces reliable structured output on one provider produces subtly malformed output on another, and the malformation is the expensive kind because it passes tests on the happy path and fails on the long tail in production.

The second is evaluation suites calibrated to the incumbent's behavior. The team that does AI seriously has built an eval harness, and that harness encodes, often unconsciously, the failure modes of the model it was built against. The thresholds, the graded examples, the regression cases are all anchored to how the current model fails. Pointed at a new model, the suite produces a wall of red that is mostly recalibration noise rather than genuine regression, and separating the two is weeks of expert judgment that the migration plan did not budget for.

The third is fine-tunes and cached context. Any investment in a provider-specific fine-tune is, by construction, non-portable: the weights live with the provider and do not come with you. Cached context is subtler. Workloads engineered around a provider's prompt-caching primitive, the kind of architecture examined in the broader treatment of the hidden cost of LLM inference and token pricing, have their economics tuned to that specific caching model. Move providers and the cache-hit economics that made the workload affordable may simply not exist in the same shape, which means the new provider can be more expensive at the workload level even when its headline token price is lower.

The fourth is safety-filter and refusal behavior. Each provider draws its content boundaries differently, and a production application quietly encodes the incumbent's boundaries into its own behavior: the prompts that route around false refusals, the post-processing that handles a particular provider's hedging, the user-facing copy calibrated to what the model will and will not say. A new provider refuses different things, hedges differently, and formats safety responses differently, and the application's accumulated accommodations are now wrong in ways that surface as user-visible regressions.

The fifth is the one that matters most at the negotiating table and is the least technical: pricing leverage that erodes the moment the provider knows you cannot credibly leave. Every renewal is a negotiation, and the buyer's only real leverage is a credible alternative. An organization that has let the first four layers accrete without managing them has, in effect, disarmed itself, and the provider's account team can read the depth of the integration as well as the buyer can, often better.

Exit drill

This is a Premium Article

Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.

Get Premium Access