| Dimension | Score |
|---|---|
| Output reliability | 3/5 |
| Workflow fit | 4/5 |
| Context handling | 3/5 |
| Learning curve vs. payoff | 4/5 |
| Failure transparency | 3/5 |
| Overall | 3.4/5 |
What AI does well
Ticket triage and routing. Zendesk AI and Intercom Fin tag intent, sentiment, and urgency on inbound tickets before a human ever opens them. I watched a 12-agent team cut average time-to-first-touch from 40 minutes to under 10. The caveat: mixed-intent tickets (“cancel my order and also fix my login”) still get tagged with only the first intent, so the second issue sits ignored.
Macro and reply drafting. Claude and ChatGPT draft full replies from a ticket plus a knowledge base article, and agents edit instead of writing from scratch. On repetitive tickets, that is a real time save, often 60 to 70 percent faster per reply. The catch: tone drifts on edge cases. Claude tends to over-apologize on billing disputes, which invites more complaints, not fewer.
Deflection on known issues. Intercom Fin resolves password resets, shipping status, and plan questions without a human, and it does this reliably because the answer space is narrow and documented. Deflection rates of 30 to 50 percent on these categories are common. Outside that narrow band, deflection accuracy drops fast.
Tagging and reporting. Zendesk AI auto-categorizes closed tickets into themes, which used to be a manual weekly task for one team lead. That freed up roughly a day a week. It still miscounts duplicate tickets from the same customer as separate issues, which inflates volume reports if nobody checks.
Where it fails
Refund and account security requests. I tested Intercom Fin on a batch of 50 refund tickets with policy exceptions (partial refund, goodwill credit, fraud flag). It approved a full refund on a ticket that should have triggered a fraud review, because the ticket language matched the “happy path” refund template closely enough. Nobody caught it until the chargeback showed up three weeks later. This is the category to keep AI out of, full stop.
Multi-ticket context across one customer. None of Intercom Fin, Zendesk AI, Claude, or ChatGPT reliably remembers that the same customer opened three related tickets last month. Each tool answers the current ticket in isolation unless someone manually pastes the history in. For a customer with a recurring hardware fault, this means three agents give three different answers before someone notices the pattern.
Escalation judgment. AI drafts a de-escalation reply for an angry customer using the same calm tone whether the customer is mildly annoyed or threatening to cancel a $50,000 contract. It does not weigh account value or churn risk unless that data is fed in explicitly, which most setups do not do by default.
My take
My take (August 2026, Bernat Sampera)
AI earns its keep on volume, not on judgment. Intercom Fin and Zendesk AI genuinely cut first-response time and handle the boring 60 percent of tickets well: resets, status checks, plan questions. That is real, measurable value, and I would not run a support team without it now.
Where I push back is on trust. The refund fraud miss I saw was not a fluke, it was a template match doing pattern recognition, not judgment. If comparing your two main drafting tools, see this ChatGPT vs Claude comparison before picking one for reply drafting, since tone and refusal behavior differ enough to matter on disputes.
Score it a 3.4. Good enough to deploy on narrow, well-documented categories. Not good enough to leave unsupervised on money or security.
What I would build
The unsolved problem is per-customer memory across tickets, agents, and tools. Right now, every AI reply generator treats each ticket as a fresh conversation, so a customer’s third complaint about the same defect gets answered like it is the first. A real fix would maintain a living case file per customer: past tickets, resolutions offered, refund history, and account value, automatically attached to every new ticket regardless of which tool or agent handles it. It would also flag when a new ticket matches a pattern from a prior unresolved issue, instead of letting five different tools each start from zero. Nobody has solved this well yet because it means unifying context across ticketing systems, chat tools, and AI drafting tools that were not built to share state.
Verdict (August 2026, Bernat Sampera): AI drafts and triages tickets fast enough to cut first response time in half, but it still mishandles or wrongly closes at least 1 in 10 refund and account access requests without flagging the miss. Overall: 3.4/5.