← AI Agents & Agentic SaaS — Blog Series
Use cases
AI Agents in Customer Support: Real Use Cases & Where to Start
The short version:
Support is the flagship agent use case — and the clearest split between deflection (kept out of the queue) and resolution (actually solved). Lead with the numbers, because support lives and dies by them.
Real benchmarks: Klarna's assistant did the work of 700 agents in month one1; Reddit cut resolution time 8.9 → 1.4 min (−84%) on Agentforce2; Intercom Fin publishes 65–76% autonomous resolution at $0.99 each3. The cautionary half: Klarna later rehired humans after an over-aggressive cutover — so this post is as much about the guardrails as the gains.
AI agents move support from a queue humans work ticket by ticket to a system that resolves the routine cases instantly and hands people only the ones that need judgment.
Why customer support is a natural fit for AI agents
Support is the flagship agent use case: high volume, repeatable patterns, and clear tools to act on (order systems, knowledge bases, account records). An agent can read a ticket, look up the relevant data, resolve it or draft a reply, and escalate the exceptions — so response times drop and human agents spend their time where empathy and judgment actually matter.
It helps to be precise about what an AI agent actually is, because the term gets stretched. An agent is not a chatbot that answers a question and stops, and it is not a rigid script that breaks the moment reality changes. An agent is given a goal in plain language, and it works through the steps to reach that goal on its own: it plans, it uses tools to take real action, it reads the result of each action, and it adjusts. That reason–act–observe loop is what lets it handle the messy, multi-step reality of customer support work — where the answer often depends on data spread across several systems and no two cases are exactly alike.
The state of AI agents in customer support (2026)
Customer support has become the flagship use case for enterprise AI agents — and the clearest example of the shift from chatbots to agents. The distinction matters: a chatbot answers a question and hopes you go away (deflection); an agent takes real actions in real systems to actually close the issue (resolution). Modern support agents interpret intent, ask clarifying questions, read and write to order and account systems, and run multi-step workflows across chat, email, and voice — escalating to a human only when they hit a genuine limit.
The 2025–2026 wave is defined by autonomous resolution platforms rather than FAQ bots: Intercom Fin, Sierra (co-founded by Bret Taylor), Decagon, Zendesk's AI agents, and Salesforce Agentforce. Vendors publish resolution figures that should be read as marketing claims: Intercom reports Fin averaging 76% resolution across 12,000+ customers (vendor claim); Decagon cites named customers such as Chime at 70% chat-and-voice resolution (vendor claim).
Pricing shifted with the technology. Instead of per-seat licences, the leaders sell outcome-based pricing — you pay per resolution or per conversation. Salesforce Agentforce publishes a list price of $2.00 per conversation; Sierra and Intercom Fin position billing around outcomes; Zendesk bills AI on a resolution allowance. The most-cited real-world deployment remains Klarna's OpenAI-powered assistant, which the company said in 2024 did the work of 700 agents and handled 2.3 million conversations in a month (vendor claims). Notably, in 2025 CEO Sebastian Siemiatkowski publicly acknowledged the aggressive cutover produced "lower quality" and said Klarna would rehire human agents — a widely reported cautionary tale about over-automation.
Customer support use cases in depth
| Use case | What the agent does | Real example |
|---|---|---|
| Autonomous ticket resolution | Interprets intent, asks follow-ups, resolves end-to-end without a human | Intercom Fin; Decagon "AI concierge" |
| Triage & routing | Classifies, tags, prioritises, routes with context | Forethought Triage Agent |
| Order/account actions | Reads/writes external systems: refunds, address & subscription changes | Fin + Shopify order status; Agentforce order management |
| Agent-assist / suggested replies | Real-time coaching and drafted responses for human agents | Cresta Agent Assist; Zendesk Copilot |
| Multilingual support | Responds across languages a human agent doesn't speak | Unbabel; Zendesk agents (80+ languages) |
| Proactive support | Detects sentiment spikes/issues and reaches out first | Zendesk sentiment analysis |
| QA & conversation analytics | Auto-scores 100% of conversations vs manual sampling | Zendesk QA (ex-Klaus); MaestroQA |
| KB gap detection & drafting | Clusters tickets to find missing articles and drafts them | HubSpot Breeze KB Agent |
The center of gravity is autonomous resolution: the agent doesn't just retrieve an answer, it acts. Fin, integrated with Shopify, can pull a customer's live delivery status and process a refund; Agentforce lists order management as a first-class capability. This is what separates 2026 agents from 2020-era bots — they hold credentials and permissions, so a "where's my order / actually, change my address / and refund the late one" conversation resolves in one thread.
Agent-assist is the quieter but faster-adopted pattern: rather than replacing the human, tools like Cresta and Zendesk Copilot analyse the live conversation and surface next steps and draft replies, compressing handle time while keeping a person accountable. Analytics use cases scale coverage — AI QA scores every interaction instead of the 1–2% a manual team samples, flagging churn risk and knowledge-base gaps.
Underpinning all of it is the pricing shift. When you pay per resolved contact rather than per seat, the vendor's incentive aligns with actually solving problems — but it also makes the resolution-versus-deflection distinction financially load-bearing, because a bot that "deflects" an abandoned chat can look successful while quietly failing the customer.
The agentic support stack
- Autonomous resolution platforms — front-line agents that close tickets: Intercom Fin, Sierra, Decagon, Zendesk AI agents, Salesforce Agentforce, Ada. These carry the outcome-based pricing and the headline resolution claims.
- Agent-assist copilots — human-in-the-loop layers that coach and draft: Cresta, Forethought Copilot, Assembled. Lower risk, immediate handle-time gains, and often the on-ramp before full automation.
- Voice AI — purpose-built for phone: PolyAI, Parloa, Cognigy (acquired by NICE), and Replicant. Voice is where 2025 consolidation was most visible.
The benefits of AI agents in customer support
The value of a well-scoped customer support agent shows up in four ways, and it compounds as the agent takes on more of the routine load:
- Speed. Work that used to sit in a queue for hours or days gets handled in seconds. The agent doesn't sleep, doesn't context-switch, and doesn't wait for the next available person.
- Consistency. The same process runs the same way every time. There's no drift between a Monday-morning task and a Friday-afternoon one, and every step is logged.
- Scale without linear headcount. Volume can double without the team doubling. People move up the value chain — from doing the task to directing and reviewing the agent that does it.
- Better use of human time. The repetitive 80% is handled automatically, so the team's attention goes to the judgment calls, the exceptions, and the relationships that actually need a person.
Support-specific pitfalls
- Hallucinated policy and legal liability. In Moffatt v. Air Canada (2024 BCCRT 149), Air Canada's chatbot invented a retroactive bereavement-fare policy; the tribunal rejected the argument that the bot was a separate entity, found negligent misrepresentation, and held the airline liable. In April 2025, Cursor's "Sam" bot fabricated a non-existent login policy and triggered cancellations. Mitigation: ground answers in a governed knowledge base (RAG), constrain policy statements to approved sources, and log every claim.
- Escalation design. A bad handoff makes customers re-explain everything. Mitigation: define explicit triggers (policy exceptions, need for judgment, any explicit human request → immediate handoff) and pass full context in a warm handoff so no re-verification is needed.
- Knowledge-base dependency. Agents are only as good as the docs behind them; stale or thin content yields confidently wrong answers. Mitigation: use gap-detection agents to flag and draft missing articles, and audit the corpus continuously.
- Tone, empathy and angry customers. Research on artificial empathy is mixed — over-emoting can feel worse, and escalated emotion is where automation struggles most. Mitigation: make "upset customer" a mandatory escalation category rather than something the bot tries to soothe.
- Over-automation frustration. Surveys repeatedly show many consumers prefer a human option; Klarna's public reversal is the cautionary example. Mitigation: always offer a visible path to a person.
How to measure a support agent
The single most important distinction is deflection versus resolution. Deflection counts contacts kept out of the human queue — a cost-avoidance number that also captures abandoned chats and wrong answers. Resolution counts problems actually solved. They diverge because a customer who gets a wrong answer and gives up registers as a "successful deflection." The guardrail: never report deflection without a re-contact rate (e.g. within 48 hours) — the one signal an over-aggressive bot cannot inflate.
- Automated resolution rate — share the agent closes end-to-end, validated against reopen/re-contact rate.
- CSAT — post-interaction satisfaction, segmented AI-handled vs human-handled.
- First-contact resolution (FCR) — solved in one interaction, no repeat.
- Average handle time (AHT) — where agent-assist shows fastest gains.
- Escalation rate — rising rates flag KB or product gaps.
- Cost per contact — Gartner figures cited by third parties put self-service near $1.84 vs assisted near $13.50.
What good looks like: two scenarios
Ecommerce order fix. A shopper messages "my order's late and the address is wrong." The agent authenticates, pulls live delivery status from Shopify, updates the shipping address, issues a goodwill refund within policy limits, and confirms — one thread, no human. Because billing is per resolution, the vendor only earns when the reopen rate stays low, so success is measured by the customer not coming back, not by the chat ending.
Emotional edge case done right. A customer contacts an airline about a bereavement fare. Rather than improvising policy (the Air Canada failure mode), the agent cites only the governed knowledge base, recognises grief as an emotional trigger, and performs a warm handoff to a human with full context attached — the customer explains nothing twice. The measurable win here isn't a lower AHT; it's a correct answer, a clean escalation, and no liability.
What to automate first
The best starting point is a task that is repeatable, rules-light, and multi-step — frequent enough to matter, but bounded enough to keep a human in the loop while you build trust. In customer support, that usually means:
- Ticket triage and routing.
- Tier-1 FAQ and knowledge-base answers.
- Suggested replies for human agents.
Pick one of these, run it in draft-and-approve mode for a couple of weeks, measure it against your baseline, and only then widen the agent's remit. This crawl-walk-run path is how teams get real value without betting the process on day one.
A practical 90-day rollout
You don't need a moonshot program to get value from a customer support agent. A focused quarter is usually enough to go from idea to a workflow the team trusts:
- Days 1–30 — pick and scope. Choose one high-volume, well-understood workflow. Write down exactly what "done well" looks like, which tools and data the agent needs, and what it is not allowed to do. Capture a baseline of today's cycle time and volume so you can prove the improvement later.
- Days 31–60 — run in draft-and-approve. Put the agent live but keep a human approving every consequential action. This is where you tune context, fix the edge cases it surfaces, and build the team's confidence. Track the human touch rate week over week.
- Days 61–90 — expand autonomy. For the cases the agent has handled cleanly and repeatably, let it act on its own while continuing to escalate the exceptions. Add the next adjacent task, and start the loop again.
By the end of the quarter you typically have one customer support workflow the agent owns end to end, a clear measure of the time it saved, and a repeatable playbook for the next one. Momentum comes from stacking small, proven wins — not from trying to automate everything at once.
Building customer support agents for your own product or team? The hard part isn't the model — it's the tool access, memory, orchestration, guardrails, and per-customer metering underneath. That's the layer Aramb handles: you describe what the agent should do in plain language and it runs reliably, with memory and live tools, reporting back what it did. It's the same foundation behind Potts, an AI coworker that lives in Slack and answers with the whole company's context — a sign that a support-grade agent can be days of work rather than months.
Customer support FAQ
What's the difference between deflection and resolution?
Deflection counts contacts kept out of the human queue — a cost number that also captures abandoned chats and wrong answers. Resolution counts problems actually solved. They diverge because a customer who gets a wrong answer and gives up looks like a "successful deflection." Never report deflection without a re-contact rate, the one signal an over-aggressive bot can't inflate.
Will an AI support agent hurt CSAT?
It can if you over-automate. Klarna's public reversal — rehiring humans after quality dropped — is the cautionary tale. The fix is to always offer a visible path to a person, make "upset customer" a mandatory escalation, and segment CSAT by AI-handled vs human-handled so you catch problems early.
What resolution rate is realistic?
Read vendor claims carefully. Intercom Fin publishes fleet-wide averages of 65–76%; some named deployments cite ~70–80%. Your real number depends on knowledge-base quality and how narrow the scope is — validate any claim against your own reopen rate.
Can the agent take real actions, or just answer?
Modern agents act: authenticated, they pull live order status, update addresses, and issue refunds within policy — resolving a multi-part request in one thread. That's the line between a 2026 agent and a 2020 FAQ bot: it holds credentials and permissions, gated for anything irreversible.
What should we automate first in support?
Ticket triage and routing, tier-1 FAQ answers, and suggested replies for human agents. Run them in draft-and-approve for a couple of weeks, measure against baseline, then widen the remit.
References & further reading
- Klarna & OpenAI — AI assistant results (Feb 2024); 2025 reversal widely reported — klarna.com/press
- Salesforce — Agentforce customer results: Reddit, City of Kyle (2025–2026) — salesforce.com/agentforce
- Intercom — Fin resolution-rate data & outcome pricing (2025–2026) — intercom.com/fin
Related: what is an AI agent? · what is agentic SaaS? · AI agents in sales
Part of an educational series on AI agents and agentic SaaS. Want this tailored to your industry or turned into a shorter version? Just ask.