TL;DR

Fresh figures this month put a number on something I've been watching happen for a year: most AI-agent projects don't fail because the technology's bad. They fail because nobody set them up to succeed. Gartner now expects more than 40% of agentic AI projects to be scrapped by the end of 2027, blaming unclear value, rising costs and weak controls. Meanwhile 72% of firms say they've got agents in production but 60% have no formal governance around them. Almost none of that is the technology's fault. It's what happens when agents get built without the experience to do it properly for the specific business they're running in. The good news for a small business is that the things that kill these projects are cheap to get right, and you don't need an enterprise budget to do it. Here's the checklist I run before I let any agent near a client's business.

There's a particular kind of story I keep hearing at the moment. A business gets excited about AI agents, builds or buys one, switches it on, and for a few weeks it's brilliant. Then something goes quietly wrong: a number drifts, a message goes to the wrong person, a task the agent was supposed to handle turns out to have been half-handled for a month. And the whole thing gets switched off in a panic, tarred as "not ready." The model was fine. The setup was the problem, and nobody had built the parts that catch a problem before it becomes an incident.

The analysts have now put figures on this. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, blaming escalating costs, unclear business value and weak risk controls. Set that next to the adoption data doing the rounds this month (72% of firms say they've got agents running for real, but 60% admit they have no formal governance around them) and the picture is clear. A lot of people have put agents to work with no plan for what happens when one goes sideways.

Why this is good news, not bad

It sounds like a warning to stay away. I read it the opposite way. The failures aren't mysterious, and they're not about the AI being too weak. They cluster around three dull, fixable causes: no clear definition of what success looks like, no access to the right data, and no plan for the day it misbehaves. Every one of those is a thing you decide, not a thing you buy. And every one comes down to the same thing: whether whoever built the agent had done it before, and understood the business well enough to get those calls right. The technology is off-the-shelf now. The judgement to apply it to your specific situation is not. Which means a small business, moving carefully, can dodge the exact traps that are sinking far bigger projects with far bigger budgets. The enterprise version fails because a committee bought a grand vision. The small-business version succeeds because one person scoped a boring job properly.

Gartner even has a phrase for part of the problem: "agent washing," where chatbots and simple scripts get relabelled as autonomous agents, then judged as if they were, and found wanting. That's worth knowing as a buyer. But it cuts deeper than labels. A lot of agents underdeliver because whoever built them hadn't built one before, and didn't yet know what separates a real agent from a script in a costume. Knowing that difference, and building for it, is most of the job.

The checklist I run before any agent goes live

This is the list I work through with a client before I'll let an agent touch anything that matters. None of it is technical. All of it is the stuff the cancelled 40% skipped.

  1. What does "working" mean, in a number? Before anything switches on, I want one sentence that says what good looks like: "flags 95% of overdue invoices within a day," not "helps with invoices." If you can't measure it, you can't tell whether it's succeeding or quietly rotting. Most dead projects never had this sentence.
  2. Does it have the data it needs, and only that? An agent starved of the right information will guess, and a guessing agent is worse than none. Equally, an agent with access to everything is a security problem waiting to happen. Give it exactly what the job requires and nothing more.
  3. What's the blast radius when it's wrong? Not if. When. If the worst case is "a slightly-off draft a human bins," proceed. If the worst case is "money left the account" or "a customer got an email you didn't see," that job doesn't run unattended until you've put a human on the trigger.
  4. Would you notice it going wrong? This is the one that catches people. If an agent could drift for three weeks before anyone spotted it, you don't have an agent, you have an unexploded problem. There has to be a check (a digest, a spot-review, a number someone watches) that surfaces trouble early.
  5. Who owns it? Not who built it. Who's responsible for it now. An agent with no owner is the one that gets switched off after the incident, because there was never anyone whose job it was to keep it healthy.
The one that does the most work: "Would you notice if it went quietly wrong?" Almost every decommissioned agent I've heard about failed this test. It wasn't that the agent did something dramatic. It's that it did something small and wrong, repeatedly, in a place nobody was looking, until the errors piled up into an incident. A five-minute weekly glance at the right number would have caught it in week one.

What this looks like when it goes right

I run my own consultancy on agents that follow exactly these rules, and that discipline is why they haven't had to be switched off in a panic. Each one has a job I could describe in a sentence, access to only what it needs, a hard line it won't cross without me, and a digest that lands in front of me so I'd spot a problem in days, not weeks. None of that is clever engineering. It's judgement, applied to the specific business each agent runs in, by someone who has built enough of them to know where they go wrong. That's the part you can't buy off a shelf. A client in property runs an agent that drafts responses to overnight enquiries and then queues them for a person to release in the morning. The drafting is automated, the sending isn't, and that single boundary is why it's been running happily for months rather than becoming a cautionary tale.

The takeaway

The headline number, four in ten projects scrapped, reads like a reason to wait. It's a map of exactly where the potholes are. The projects that die don't die from weak models; they die from skipped basics, and the basics are free. Knowing which ones matter for your business is where experience earns its place. If you're thinking about putting an agent to work, don't start by choosing the technology. Start with the specific job and the five questions above, ideally with someone who has built agents before and understands how your business actually runs. If you can answer all five cleanly, you're already in the 60% that survives. If you can't, you've just found the work to do first. And it's a lot cheaper to do it now than to explain an incident later.