Cutting Headcount Is Not an AI Customer Support Strategy
In February 2024, Klarna said its AI assistant had handled 2.3 million conversations in a single month — two-thirds of all its chats — at an average resolution time under two minutes, doing the work of 700 agents. Fifteen months later the company was recruiting humans again.
The bot hadn't got worse. The number the AI customer support strategy was built around was the wrong number.
Klarna's CEO was unusually direct about why. "As cost unfortunately seems to have been a too predominant evaluation factor when organizing this, what you end up having is lower quality," Sebastian Siemiatkowski told Bloomberg in May 2025, as reported by CX Dive. Not "the model hallucinated." Not "we chose the wrong vendor." Cost was the evaluation factor. Quality was the thing that gave way.
We've sat in a lot of evaluation calls for ChatterMate, and the first question is almost always some version of how many agents does this replace. Fair question. It's also the single best predictor we've found of a deployment that goes badly.
Headcount is the only number everyone in the room agrees on
An agent has a salary, a manager, a seat licence, an onboarding cost and an attrition rate. Put all of that in a spreadsheet and it sits there quietly, not arguing with anyone. Finance likes it. The board understands it. You can forecast it three years out.
Resolution quality has none of those properties. What is a customer worth who got a trustworthy answer at 11pm on a Sunday? What does it cost you when someone gets a confidently invented refund policy and repeats it to nine other people? Nobody has a clean figure. So nobody puts it in the model.
The business case then gets written in the only currency that's legible, and everything else gets demoted to a line that reads "monitor CSAT." Which is how you end up with a project whose success condition is "fewer people" rather than "fewer angry people."
Two very different systems clear those two bars. And an AI customer support strategy inherits whichever bar you wrote down first.
What the ROI data actually shows
Gartner analysed 432 AI use cases in customer service and found that only about a quarter produce a return on investment. Another quarter deliver negative returns. Eleven percent break even. And 42% — the largest group by some distance — sit in a bucket where support leaders say they simply don't know what value was produced (CX Dive, August 2026).
Read that middle number again. A quarter of these projects are actively losing money for the companies running them.
Meanwhile more than three-quarters of leaders plan to increase AI spend this year, teams are running nearly five use cases each, and roughly 13% of the function's budget is going to AI. Gartner also found that 56% of service leaders expect their own incentives to be tied to AI outcomes in 2026 — outcomes that, per the same research, 42% of them can't measure.
That is a lot of career risk stacked on top of a metric nobody has defined.
Why can't they measure it? Because the instrumentation was never built. If the project's success condition was "reduce headcount," then the only thing anyone wired up was a ticket counter and a monthly invoice. There's no baseline resolution rate from before the rollout, no sample of conversations anyone graded by hand, no record of which answers were grounded in a real document and which were improvised. Twelve months later someone asks what it's worth and the honest answer is a shrug.
Antoine Nasr, who leads AI at Forethought AI Agents by Zendesk, described the pattern to CX Dive as top-down: "everyone is trying to get AI." Geller, from Info-Tech, said the same thing more sharply — too many rollouts start with pressure to show the board a credible AI story rather than with a defined customer problem. Start there and the metric you end up with is whatever was easiest to count.
Why an AI customer support strategy that starts with headcount fails
Three specific failure modes, all of which follow from the same starting error.
Containment gets mistaken for resolution
Julie Geller, principal research director at Info-Tech Research Group, put it about as well as it can be put: "Delaying contact with a human agent is not the same as resolving the customer's problem. The real test is much simpler: did the customer get what they needed, with less effort?"
A bot optimised for containment learns to be sticky. It asks clarifying questions it doesn't need. It offers a help-centre article instead of an answer. It buries the handoff. Every one of those behaviours improves the dashboard and worsens the experience, which is exactly why containment rate is such a treacherous number to manage against.
The customer's version of the metric is different and much harder to game: did this end?
The saving comes back wearing a different job title
Here's the part that should worry any CFO signing off on a headcount-driven case. Gartner found that the share of customer service organisations growing headcount is roughly the same as the share shrinking it — about a quarter each way. Teams that deploy AI also tend to hire new specialist roles to run it: people to curate the knowledge base, review escalations, tune the retrieval, own the evals.
Gartner's own forecast is blunter still. By 2027, half of the organisations that expected to significantly reduce their customer service workforce will abandon those plans. In a March 2025 poll of 163 service leaders, 95% said they intend to keep human agents in the loop.
"While AI offers significant potential to transform customer service, it is not a panacea," said Kathy Ross, the Gartner analyst behind that prediction. "The human touch remains irreplaceable in many interactions."
So the honest version of the cost case isn't salary out, subscription in. It's some salary out, subscription in, some new salary in, and a permanent maintenance obligation you didn't previously have. Still often worth it. Just not the number on the slide.
The escape hatch stops being a priority
When the goal is deflection, the path to a human is a leak. When the goal is resolution, it's a feature — and one that most teams under-build. A bot that says "let me get someone" and then drops the customer into a fresh queue with none of the conversation attached has not helped anybody; it has added a step. We wrote about what a handoff has to carry to not be infuriating because we kept seeing the same broken pattern.
Klarna's reversal was, at bottom, about this. Siemiatkowski's line to Bloomberg was that customers should always have a human "if you want."
The strongest version of the opposing case
Now the other side, properly, because there is one and it isn't stupid.
Forrester expects half of today's customer service jobs to be gone to AI by 2030. Their analyst Max Ball's framing — that a lot of this work never required human-level intelligence in the first place — is hard to dismiss if you've ever read a support queue at 2am. Password resets. Order status. "Where is my invoice." That work is genuinely low-value for a person to do and genuinely miserable to do all day.
And the cost pressure is not imaginary. Support has always been a cost centre with a CFO looking at it. Pretending otherwise is how vendors sell software to people who don't have to pay for it.
Klarna's own numbers make the point better than we can. Even after the reversal, the bot still handles two-thirds of all inquiries. Response times are up 82% on where they were. Repeat issues fell 25%. That is real, durable value that no amount of rehiring undoes. The company didn't conclude AI failed. It concluded that cost had been weighted too heavily in the design, and rebalanced.
Which is the argument, restated: automation of the repetitive stuff is correct, and saving money is a legitimate outcome. It's just a terrible design target. Aim at resolution and you often get the saving. Aim at the saving and you frequently get neither.
What we'd put in the spreadsheet instead
None of the below is harder to measure than headcount. It's just less familiar, so it takes one quarter of discipline to get going.
Build an eval set before you buy anything. Pull 100 real tickets from the last month, spread across your actual mix — not the easy ones. Write down the correct answer for each. That set is now your benchmark, and it will tell you more in an afternoon than any vendor demo. Run it again after every knowledge-base change.
We've watched this exercise change people's minds in both directions, which is how you know it's a real test. One team came in expecting a disaster and found the bot handled 71 of their 100 cleanly. Another came in confident and discovered that every question touching their returns policy produced a plausible, well-written, wrong answer — because two help-centre pages contradicted each other and had done for a year. No model was going to fix that. A twenty-minute edit did.
Measure resolution, not containment. A conversation counts as resolved when the customer didn't come back about the same thing within seven days. That's it. It's a slightly annoying query to write and it is worth every minute.
Track answer provenance. What share of the bot's answers cite a specific source document? Answers with no source behind them are the ones that invent policy. Grounding every reply in your own docs, with the citation visible, is the whole reason we built ChatterMate the way we did — and it makes the failure mode auditable instead of mysterious.
Read your escalations weekly. Not the count. The reasons. Escalation transcripts are the highest-signal document your support function produces, and roughly nobody reads them.
Count the tickets that never arrived. When the bot repeatedly fails on a question, that's usually a documentation gap, not a model problem. Fix the doc and the volume drops at the source — for the bot and the humans and the search engine, all at once.
Then, at the end of the quarter, look at cost. It'll be there. It just shouldn't be the thing steering.
The uncomfortable part
If your AI project only makes sense when you assume the layoff, you don't have an AI customer support strategy. You have a redundancy plan with a chat widget attached, and the market has now run that experiment publicly enough that we can all read the results.
The teams doing this well are boring about it. They pick a narrow set of questions, ground the answers in real documentation, watch the escalations, and expand slowly. Their deflection numbers go up more slowly than the case study promised. Their customers stay.
Klarna spent fifteen months and a fair amount of brand equity learning the same thing. That lesson is free now. Take it.
Written by the ChatterMate team — we build an open-source AI support agent that answers from your own docs and cites what it used. More of our thinking on this in per-seat pricing is quietly dying and agent washing, or browse the full blog.
If you'd rather test this than argue about it: ChatterMate is open source and free to start — your first 300 chats cost nothing, and you can self-host the whole thing if you'd prefer your support transcripts stayed on your own infrastructure. Point it at your docs, run your 100 tickets through it, and see what the resolution number actually says.

Aug 26,2026
By runix