AI Containment Rate: Why 76% and 14% Are Both True
Two numbers, published four months apart, both from credible sources, both describing what people loosely call an AI containment rate.
Intercom says Fin's average resolution rate across more than 7,000 teams now sits at 76%. Gartner surveyed 5,728 customers and found that only 14% of their service issues were fully resolved in self-service — and for issues customers themselves described as "very simple," the figure was still just 36%.
Sixty-two points apart, on what sounds like the same question. Neither party is lying. They're counting different things, from different ends of the interaction, and the difference is the single most useful thing a support team can understand before it signs an AI contract.
What an AI containment rate actually counts
Containment is a vendor-side metric. It measures the share of conversations that started in the bot and ended in the bot — no human ever touched the thread.
That's it. That's the whole definition.
Notice what it does not measure. It doesn't ask whether the customer got what they came for. It doesn't ask whether they gave up and emailed sales instead. It doesn't ask whether they came back three days later with the same question, opening a fresh conversation that counts as its own contained interaction. A customer who types their question, reads a confident wrong answer, closes the tab, and switches to a competitor is a perfectly contained conversation.
Gartner's 14% comes from the other end. They asked the human being whether the issue was resolved. Different denominator, different judge, different answer.
Both numbers are honest inside their own frame. The trouble starts when a buyer reads a 76% on a pricing page and mentally converts it into "76% of my customers will go away happy." Those are not the same claim, and no vendor benchmark — ours included — can make the second one for you.
The gap is mostly the denominator
Three things drive most of the divergence, and none of them require anyone to be acting in bad faith.
Vendor averages are weighted by the customers who deploy well. A company with a maintained, structured help center and a narrow product surface will land far above the mean. A company whose docs were last touched in 2023 will land far below it. The published average is a blend, and it tells you almost nothing about which end you're on. Your knowledge base quality is the biggest single variable, and it's entirely in your hands.
Ticket mix decides the ceiling before the model does. "Where's my order" and "how do I reset my password" are trivially automatable. Billing disputes, cancellations with retention implications, and anything with an angry human on the other end are not — not because the model can't parse them, but because you probably shouldn't want them fully automated. A store with 60% WISMO traffic and a B2B platform with 60% configuration questions will get wildly different containment from identical software.
Customers are measured on the whole journey; bots are measured on one session. Gartner's respondents were describing an issue, which might span a help-center search, a chat, an email, and a phone call. Containment scores the chat in isolation. When someone bounces between four channels, three of those touchpoints can be technically contained while the underlying issue stays open.
Read a benchmark, then, as a range rather than a promise. Independent testing of the major AI agents tends to land well below the headline figures, and the sensible planning assumption for a first deployment sits somewhere in the 40–60% band depending on your mix — with the honest caveat that your own first month of data will beat any industry number as a predictor.
The industry quietly changed the definition this year
Here's the part that got surprisingly little attention.
In March, Intercom announced it was moving its pricing metric from resolutions to outcomes. The reasoning is sound and worth reading in full: as Fin took on multi-step work — gathering context, reading and writing to external systems, executing actions — a lot of genuinely valuable work ended with a deliberate handoff to a human. Under a resolution-only metric, all of that counted as a failure. So they widened the definition to include it.
Credit where it's due. That's the right call for customers, and it took some nerve to make it while resolutions were the number everyone quoted.
But look at what it means for benchmarking. The most-cited resolution rate in the industry now sits on a definition that shifted mid-flight. Anyone comparing a 2025 number to a 2026 number is comparing two different measurements with the same name. And every vendor that follows — most will — will do it in their own way, with their own boundary between "outcome" and "handoff."
We think this is a decent argument for treating every vendor-reported containment or resolution figure as marketing copy rather than data. Not dishonest. Just not comparable, and not yours.
Optimising for containment eventually hurts you
There's a failure mode that shows up once containment becomes a target rather than an observation. Teams start removing the escape hatch.
Bury the "talk to a human" button. Require two failed AI attempts before escalation is even offered. Set the confidence threshold low enough that the bot always says something. Every one of those moves the number up. Every one of them is a bad idea.
Gartner's August 2026 survey of 3,566 customers puts a figure on the risk: 87% say it's essential to have the option to reach a human agent when a company uses generative AI. Only 50% said the AI made things easier. Eric Keller, the Gartner analyst on the study, was direct about it: "Service leaders should not use GenAI as a mandatory first step for every issue." When customers get pushed through repeated failed AI attempts before they can reach a person, they stop using the tool at all.
There's a second-order effect too. That same Gartner research found customers were roughly three times more likely to use a third-party tool like ChatGPT or Gemini for their most recent service interaction than a company-provided chatbot. If your bot is a wall, people route around it — and you lose the interaction data along with the goodwill.
The deeper problem is that a wrong-but-confident answer scores identically to a right one. That's the same structural flaw behind most AI chatbot hallucinations: nothing in the metric penalises invention. Grounding answers in your actual documentation and showing the source is what closes that hole, and it's why we made citations non-optional in ChatterMate rather than a setting.
A scorecard that's harder to game
Containment still belongs on the dashboard. It just shouldn't be the headline. Four things we'd put alongside it:
Reopen rate within 7 days. If a contained conversation is followed by the same customer asking the same thing again, it wasn't resolved — it was deferred. This is the cheapest possible sanity check on a containment number and almost nobody runs it.
Escalation quality, not escalation volume. A handoff where the AI has already collected the order number, verified the account and summarised the issue is worth more than a resolution on a question nobody cared about. Measure the minutes your agents save on escalated threads. Intercom's shift to "outcomes" is essentially an admission that this category was invisible before.
CSAT split by intent, not in aggregate. Blended CSAT hides everything interesting. Structured requests — password resets, order status, plan changes — score well under automation. Complaints and billing disputes score badly, and should, because those customers want a person. If you only look at the average you'll draw exactly the wrong conclusion about where to automate next.
Coverage gaps. Gartner found the most common reason self-service failed was simple: 43% of customers couldn't find content relevant to their issue. Not a model problem. A documentation problem. Every question your AI couldn't answer from your docs is a ranked to-do list for whoever owns your help center, and working that list moves containment more reliably than any prompt tuning.
That last one is where we've seen the most consistent movement. Teams arrive expecting to tune the AI and end up rewriting six help articles instead — which is how training a chatbot on your docs actually works once you get past the demo.
Where this is heading
Gartner's widely-quoted forecast is that agentic AI will autonomously resolve 80% of common customer service issues by 2029, cutting operational costs 30%. Note the word "common" — it's doing an enormous amount of work in that sentence, and the prediction is much less dramatic once you read it as "80% of the easy half."
Meanwhile the expectations keep tightening. Zendesk's 2026 CX Trends report, drawing on more than 11,000 consumers and leaders across 22 countries, found 85% of CX leaders say customers will drop a brand that can't resolve an issue on first contact, and 95% of consumers expect a clear explanation for AI-made decisions. First-contact resolution and explainability. Neither is captured by containment.
So the metric is being squeezed from both sides. Vendors are broadening it because it undercounts real work. Customers are pushing back on the behaviours that inflate it. Somewhere in the middle is a number that's still useful for tracking your own trend line week over week and close to useless for comparing two products.
Track yours. Ignore everyone else's. And if a vendor quotes you a containment figure without telling you the denominator, ask — the answer tells you more about the company than the number does.
Written by the ChatterMate team — we build an open-source AI support agent that answers from your docs and cites its sources. More on how the economics stack up in our 2026 AI customer support trends piece, and our argument that the support ticket queue is a relic. The rest of our writing lives on the ChatterMate blog.
If you'd rather test this on your own numbers than argue about someone else's, ChatterMate is open source and free to start — your first 300 chats cost nothing, and you can self-host the whole thing if you'd rather your support data never left your infrastructure.

Aug 23,2026
By runix
AI Chatbot Testing: Build the Eval Set Before You Launch
5 day ago[…] rate is measuring how often customers gave up, not how often they got helped — which is why the same product can honestly claim 76% and 14% depending on what it counts. An eval suite is what makes a deflection number mean something, […]