AI Support Deflection Rates Are Measuring the Wrong Thing

AI Support Deflection Rates Are Measuring the Wrong Thing

Customer service is the AI deployment with the cleanest return-on-investment story in the enterprise. The cost base is large and legible, the work is text, and the vendor can show a dashboard within weeks that says most conversations no longer reach a human. The saving arrives as avoided headcount, the easiest line in the budget to believe.

The problem is the number underneath the story. Deflection, also called containment, is the share of support conversations that end without reaching a human agent. It counts every one of those conversations as a win: the customer who got a correct answer, the customer who gave up, the customer who called the next morning, and the customer who quietly moved their business elsewhere. It measures where the conversation stopped, not whether the problem did.

The thesis of this piece: the vendors who now price per resolution have pushed the industry toward a better definition, but most enterprises still manage to the old one, and they are about to discover the gap in their retention data rather than in their support dashboard. The fix is not a better chatbot. It is a better measurement stack, a contract that defines "resolved" in the buyer's terms, and a routing policy that automates where a wrong answer is cheap and hands off early where it is not.

Measurement stack

How Deflection Became the Headline Number

Deflection became the default for a mundane reason: it is the only outcome metric the support platform can compute on its own. The platform knows whether a conversation was escalated to a human queue. It does not know whether the customer's card was actually unblocked, whether the refund landed, or whether the same person rang the call centre an hour later from a different phone number. Containment is what the system can see, so containment is what the system reports.

That made sense when self-service meant a help-centre article and a search box. It makes much less sense when a generative agent handles most inbound contact and the business case assumes the contained share is also the resolved share. The evidence says the two usually diverge.

Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service issues were fully resolved in self-service, and that even for issues customers described as "very simple", only 36% resolved fully without another channel. Nearly nine in ten journeys that began in self-service ended up being resolved through multiple channels such as phone or email (Gartner, August 2024). A self-service channel reporting high containment and a customer base reporting low resolution are not contradictory. They are the same system described from two ends.

Who defines resolved

Where the Number Overcounts

The overcount comes from three specific construction choices, and each one is fixable once it is named.

Abandonment is scored as success

A conversation that ends without escalation is contained. A customer who closes the window after the third unhelpful answer has ended the conversation without escalation. The metric cannot tell them apart, and a poorly performing agent can raise its containment rate simply by making escalation harder to find.

Repeat contacts are never joined back

The customer who fails in chat does not disappear. They phone, email, post on social media, or walk into a branch. Unless the organisation joins those contacts back to the original conversation by customer identity and issue, the chat is recorded as contained and the phone call as a fresh contact. The AI channel looks efficient, the voice channel looks unexpectedly busy, and nobody connects the two because they are owned by different teams with different dashboards. The Commonwealth Bank of Australia case below turned on exactly this: the bot's reported effect on call volume and the call centre's lived experience pointed in opposite directions.

The helpfulness survey nobody answers

The usual corrective is a "was this helpful?" prompt at the end of the conversation. Response rates on these prompts are small, and the people who respond are not a random sample: satisfied customers click yes on the way out, and the most frustrated customers are the least likely to stay long enough to click anything. A satisfaction score computed on the minority who answer says little about the majority who did not.

Routing grid

This is a Premium Article

Sign up for a Premium membership to read this article and get full access to strategic intelligence on technology and business.

Get Premium Access