Guides

Why AI support projects fail — and the checks that prevent it

Gartner expects more than 40% of agentic AI projects to be cancelled by 2027.

August 15, 2026 · 8 min read
Six common causes of AI support project failure and the pre-launch check that catches each
Six common causes of AI support project failure and the pre-launch check that catches each

Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls (Gartner).

That is a striking number for a technology this widely adopted. But the failures are not exotic, and they are mostly not technical. They cluster into six causes, and every one is visible before you sign anything.

The short answer

CauseHow it shows upThe check that catches it
No defined success measureSix months in, nobody can say if it workedAgree one number, and today's baseline, before launch
Data that is not readyThe agent answers confidently and wronglyRead what it will read. Is it current? Who owns it?
Scope too broadMediocre at everything, trusted for nothingThree intents at launch. Not thirty
No escalation designCustomers trapped in loopsTest the handoff before the happy path
Nobody owns it after launchQuality decays quietly over monthsName the owner and the weekly review slot
Risk controls added lastSecurity review blocks go-live — or an incident doesDecide permissions and audit logging up front

1. Nobody defined what success looks like

"We want AI in customer service" is not an objective. It cannot be measured, so the project cannot end, so it drifts until someone cancels it.

The fix costs an hour: pick one primary number, record where it stands today, and set a review date. Candidates that work — percentage of enquiries resolved without a human, median first-response time, after-hours enquiries answered, bookings completed without staff involvement.

Worth knowing what actually moves. Salesforce found that the KPI organisations most often report improving after deploying AI agents is customer satisfaction — ahead of agent productivity, handle time and retention (Salesforce). If you justified the project purely on headcount savings, you may be measuring the wrong thing and reporting a failure that is actually a success.

2. The data was not ready

This is the most common cause and the least discussed, because it is nobody's exciting project. An AI agent answers from what you give it. Give it a price list from 2024 and it will quote 2024 prices — fluently, confidently, to customers.

The people closest to the work already know. Salesforce found 72% of service operations professionals call data readiness a major blocker, against 59% of service leaders (State of Service, 7th edition) — a gap worth noticing, because the people who will run the thing are more worried than the people approving it.

The check: before launch, read the documents the agent will read. Not a sample — the actual set. For each: is it current, who updates it, and how will the agent learn that it changed? If the answer to the last is "someone will remember", it will not happen.

3. The scope was too broad

Attempting every enquiry type at once produces a system that is adequate at all of them and trusted for none. The first bad answer in a category nobody tested becomes the story everyone repeats internally, and support for the project evaporates.

The pattern that survives contact with reality: three intents, chosen because they are high-volume and low-risk. Opening hours. Availability. Booking. Get those genuinely right, measure, then widen.

This is also the difference between organisations getting value and those stuck in pilots. McKinsey found 88% of organisations use AI in at least one function, but only around 23% are scaling an agentic system, and no more than 10% are scaling agents in any single business function (The State of AI, November 2025). Experimenting is easy. The gap between experiment and production is mostly discipline about scope.

4. Escalation was an afterthought

Every demo shows the happy path. Almost none show what happens when the agent cannot help — and that is the interaction which determines whether customers trust the channel at all.

Bad escalation looks like: repeating "I'm sorry, I didn't understand" three times; offering a human and then doing nothing; handing over with no context so the customer explains everything again.

The check: before testing whether it answers well, test whether it fails well. Ask something outside its scope. Say "I want to speak to a person". Say something ambiguous. Then ask: how many turns to reach a human, and what does that human see?

A useful rule: an agent that cannot help should reach a human within two turns, carrying the full conversation with it.

5. Nobody owned it after launch

AI support degrades quietly. Products change, policies change, documents drift, and the agent keeps confidently answering from what it knew. There is no error log for "answered an outdated question smoothly" — it just gradually gets worse until someone notices complaints.

The check: name a person and a recurring slot before go-live. Thirty minutes a week reading transcripts — especially escalations and abandonments — catches more problems than any dashboard. This is the most skipped step on this list and the cheapest.

6. Risk controls were left until last

Permissions and audit logging get treated as a pre-launch formality, then block go-live for a month — or worse, do not, and something goes wrong in production.

McKinsey found that security and risk concerns are the top barrier to scaling agentic AI, ahead of regulatory uncertainty and technical limitations, and frames the shift precisely: with agents you must contend not only with a system saying the wrong thing but doing the wrong thing — taking unintended actions, misusing tools, or operating outside its guardrails (State of AI trust in 2026).

Decide these before you build:

  1. What can it read? Table by table, field by field. Payment details and internal notes usually should not be in that set.
  2. What can it change? Every write is a way to be wrong at scale.
  3. What needs approval? Refunds, cancellations, anything past a monetary threshold.
  4. What is logged? You need to answer "what did it say to this customer, and why" months later.
  5. Where does data live, and for how long? Especially transcripts.

The pre-launch checklist

Ten questions. If you cannot answer all of them you are not ready — and answering them takes a morning, against a project that takes months.

  1. What single number are we trying to move, and what is it today?
  2. Which three enquiry types launch first, and why those?
  3. What exactly will the agent read, and who keeps it current?
  4. What can it change, and what requires human approval?
  5. How does a customer reach a human, and in how many turns?
  6. What does that human see when the handover happens?
  7. Who reads the transcripts, and when is it in their calendar?
  8. What is logged, and for how long is it kept?
  9. What is our rollback if quality is unacceptable in week one?
  10. When do we review, and what result would make us stop?

Question ten is the one people resist and the one that matters most. A project with no defined failure condition cannot be evaluated — only defended.

Frequently asked questions

Is the 40% cancellation figure a reason not to start?

No — it is a reason to scope tightly and measure honestly. The same body of research shows organisations getting value quickly when they do: Salesforce reports 70% see measurable value within 60 days. Failure clusters around vague objectives and broad scope, both of which are choices.

How long before we know if it is working?

For a narrow deployment with a defined metric, four to eight weeks is usually enough to see direction. If you cannot tell after three months, the problem is almost always the measure rather than the technology.

Should we run a pilot first?

Yes, but bound it. An open-ended pilot is how projects die of exhaustion. Define what it must show, by when, to become production.

What if our team resists it?

Usually a signal worth listening to. Support teams know which enquiries are genuinely repetitive and which need judgement — and they are right more often than the plan is. Involving them in choosing the first three intents converts the most credible sceptics into the people who make it work.

Can we start without cleaning up our data?

You can start on a narrow slice whose data you trust — one product line, one service. What you cannot do is point an agent at a decade of unmaintained documents and expect it to sort out which are current.

The takeaway

AI support projects rarely fail because the technology could not do the job. They fail because nobody agreed what the job was, the information was stale, the scope was too wide, the failure path was untested, and no one owned it afterwards.

All five are decisions, all five are cheap to get right at the start, and all five are expensive to fix once live.

Planning a deployment? Serve AI is built around exactly these controls — per-workspace data isolation, column-level limits on what the AI may read, and a full transcript of everything it ever said. See how the controls work or talk through your rollout.

Never miss another customer

Serve AI answers your calls, chats and WhatsApp messages 24/7 — in your own voice.

Start free — no card required →
← All insights