Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls (Gartner).
That is a striking number for a technology this widely adopted. But the failures are not exotic, and they are mostly not technical. They cluster into six causes, and every one is visible before you sign anything.
The short answer
| CauseHow it shows upThe check that catches it | ||
| No defined success measure | Six months in, nobody can say if it worked | Agree one number, and today's baseline, before launch |
| Data that is not ready | The agent answers confidently and wrongly | Read what it will read. Is it current? Who owns it? |
| Scope too broad | Mediocre at everything, trusted for nothing | Three intents at launch. Not thirty |
| No escalation design | Customers trapped in loops | Test the handoff before the happy path |
| Nobody owns it after launch | Quality decays quietly over months | Name the owner and the weekly review slot |
| Risk controls added last | Security review blocks go-live — or an incident does | Decide permissions and audit logging up front |
1. Nobody defined what success looks like
"We want AI in customer service" is not an objective. It cannot be measured, so the project cannot end, so it drifts until someone cancels it.
The fix costs an hour: pick one primary number, record where it stands today, and set a review date. Candidates that work — percentage of enquiries resolved without a human, median first-response time, after-hours enquiries answered, bookings completed without staff involvement.
Worth knowing what actually moves. Salesforce found that the KPI organisations most often report improving after deploying AI agents is customer satisfaction — ahead of agent productivity, handle time and retention (Salesforce). If you justified the project purely on headcount savings, you may be measuring the wrong thing and reporting a failure that is actually a success.
2. The data was not ready
This is the most common cause and the least discussed, because it is nobody's exciting project. An AI agent answers from what you give it. Give it a price list from 2024 and it will quote 2024 prices — fluently, confidently, to customers.
The people closest to the work already know. Salesforce found 72% of service operations professionals call data readiness a major blocker, against 59% of service leaders (State of Service, 7th edition) — a gap worth noticing, because the people who will run the thing are more worried than the people approving it.
The check: before launch, read the documents the agent will read. Not a sample — the actual set. For each: is it current, who updates it, and how will the agent learn that it changed? If the answer to the last is "someone will remember", it will not happen.
3. The scope was too broad
Attempting every enquiry type at once produces a system that is adequate at all of them and trusted for none. The first bad answer in a category nobody tested becomes the story everyone repeats internally, and support for the project evaporates.
The pattern that survives contact with reality: three intents, chosen because they are high-volume and low-risk. Opening hours. Availability. Booking. Get those genuinely right, measure, then widen.
This is also the difference between organisations getting value and those stuck in pilots. McKinsey found 88% of organisations use AI in at least one function, but only around 23% are scaling an agentic system, and no more than 10% are scaling agents in any single business function (The State of AI, November 2025). Experimenting is easy. The gap between experiment and production is mostly discipline about scope.
4. Escalation was an afterthought
Every demo shows the happy path. Almost none show what happens when the agent cannot help — and that is the interaction which determines whether customers trust the channel at all.
Bad escalation looks like: repeating "I'm sorry, I didn't understand" three times; offering a human and then doing nothing; handing over with no context so the customer explains everything again.
The check: before testing whether it answers well, test whether it fails well. Ask something outside its scope. Say "I want to speak to a person". Say something ambiguous. Then ask: how many turns to reach a human, and what does that human see?
A useful rule: an agent that cannot help should reach a human within two turns, carrying the full conversation with it.
5. Nobody owned it after launch
AI support degrades quietly. Products change, policies change, documents drift, and the agent keeps confidently answering from what it knew. There is no error log for "answered an outdated question smoothly" — it just gradually gets worse until someone notices complaints.
The check: name a person and a recurring slot before go-live. Thirty minutes a week reading transcripts — especially escalations and abandonments — catches more problems than any dashboard. This is the most skipped step on this list and the cheapest.
6. Risk controls were left until last
Permissions and audit logging get treated as a pre-launch formality, then block go-live for a month — or worse, do not, and something goes wrong in production.
McKinsey found that security and risk concerns are the top barrier to scaling agentic AI, ahead of regulatory uncertainty and technical limitations, and frames the shift precisely: with agents you must contend not only with a system saying the wrong thing but doing the wrong thing — taking unintended actions, misusing tools, or operating outside its guardrails (State of AI trust in 2026).
Decide these before you build:
- What can it read? Table by table, field by field. Payment details and internal notes usually should not be in that set.
- What can it change? Every write is a way to be wrong at scale.
- What needs approval? Refunds, cancellations, anything past a monetary threshold.
- What is logged? You need to answer "what did it say to this customer, and why" months later.
- Where does data live, and for how long? Especially transcripts.
The pre-launch checklist
Ten questions. If you cannot answer all of them you are not ready — and answering them takes a morning, against a project that takes months.
- What single number are we trying to move, and what is it today?
- Which three enquiry types launch first, and why those?
- What exactly will the agent read, and who keeps it current?
- What can it change, and what requires human approval?
- How does a customer reach a human, and in how many turns?
- What does that human see when the handover happens?
- Who reads the transcripts, and when is it in their calendar?
- What is logged, and for how long is it kept?
- What is our rollback if quality is unacceptable in week one?
- When do we review, and what result would make us stop?
Question ten is the one people resist and the one that matters most. A project with no defined failure condition cannot be evaluated — only defended.
Frequently asked questions
Is the 40% cancellation figure a reason not to start?
No — it is a reason to scope tightly and measure honestly. The same body of research shows organisations getting value quickly when they do: Salesforce reports 70% see measurable value within 60 days. Failure clusters around vague objectives and broad scope, both of which are choices.
How long before we know if it is working?
For a narrow deployment with a defined metric, four to eight weeks is usually enough to see direction. If you cannot tell after three months, the problem is almost always the measure rather than the technology.
Should we run a pilot first?
Yes, but bound it. An open-ended pilot is how projects die of exhaustion. Define what it must show, by when, to become production.
What if our team resists it?
Usually a signal worth listening to. Support teams know which enquiries are genuinely repetitive and which need judgement — and they are right more often than the plan is. Involving them in choosing the first three intents converts the most credible sceptics into the people who make it work.
Can we start without cleaning up our data?
You can start on a narrow slice whose data you trust — one product line, one service. What you cannot do is point an agent at a decade of unmaintained documents and expect it to sort out which are current.
The takeaway
AI support projects rarely fail because the technology could not do the job. They fail because nobody agreed what the job was, the information was stale, the scope was too wide, the failure path was untested, and no one owned it afterwards.
All five are decisions, all five are cheap to get right at the start, and all five are expensive to fix once live.
Planning a deployment? Serve AI is built around exactly these controls — per-workspace data isolation, column-level limits on what the AI may read, and a full transcript of everything it ever said. See how the controls work or talk through your rollout.