Why AI projects fail small firms, and what the ones that worked did differently
GUIDE · FOR THE SCEPTIC · 11 MIN
Facts on this page verified August 2026.
The short answer: almost never the model, almost always the setup
They fail on decisions made before any software is chosen. Four causes account for most of it. Nobody wrote down the number the project was supposed to move. Nobody owned it once it was live. The machine sat beside the real work instead of inside it. And the scope was picked because it was interesting rather than because it was bleeding. All four are organisational, all four are cheap to fix at twenty people, and none of them is a question about which model you use.
The number everyone quotes, read properly
You have seen the headline: ninety-five percent of corporate AI projects return nothing. It comes from one study, it is a real finding, and almost every article repeating it gets the scope wrong. Here is what it measured.
95% · 5%
What it counted was large organisations, not firms like yours. One hundred and fifty interviews with leaders, a survey of three hundred and fifty employees, and three hundred publicly disclosed deployments, all at corporate scale. It is a finding about generative-AI programmes inside big companies. It is not a finding that the technology does not work, and it is not a prediction that a twenty-person firm putting an agent on its phone will fail nineteen times out of twenty.
The second number is the one people attribute to the wrong publisher. Search for AI failure rates and you will be told that RAND found more than eighty percent of AI projects fail. Open the report and the sentence is in the introduction, with a footnote to a magazine commentary.
>80%
What RAND actually did is more useful than the number it borrowed. Sixty-five interviews with data scientists and engineers of at least five years standing, conducted between August and December 2023, producing five root causes. Four of the five are organisational: the wrong problem, the wrong data, chasing the technology instead of the user, and infrastructure that cannot carry a deployed model. Only the fifth is about the limits of the technology itself.
The third number says the failures are being noticed, and acted on. Abandonment climbed sharply in one year, which is a healthier signal than it sounds: firms are now stopping things that are not working instead of leaving them switched on.
42% · 17%
Read the three together and the shape is clear. Failure is common, it is rising, and it is overwhelmingly organisational. Every one of these studies looked at companies far larger than yours, which is a limit on what they prove and also the opening: the causes they name are cheaper to avoid at twenty people than at twenty thousand, because at twenty people the person who decides is the person who watches it run.
The five ways it actually goes wrong
Each mode below is what it looks like from inside the firm, followed by what the projects that worked did instead. None of them is exotic. Every one of them is decided in the first fortnight, usually in a conversation nobody wrote down.
One: there was never a before-number. The project was approved on a feeling that something was slow, and it is judged a year later on a different feeling. Nobody can say what the phone was costing before, so nobody can say what it costs now, and the line item quietly loses the next budget argument to something that can show a figure. This is the most common of the five and the cheapest to have avoided.
Two: nobody owned it after the launch. The build finished, the supplier left, and the automation became everybody's and therefore nobody's. Six weeks later an opening hour changed, or a price list moved, and the agent started giving an answer that used to be right. Nobody had the job of noticing, so the fix was to switch it off.
Three: it sat beside the work instead of inside it. The machine produced something correct and then somebody had to copy it into the system where the work actually happens. That copying is the tax that kills adoption: it is small enough that nobody complains and large enough that people stop bothering within a month. A tool that lives in a separate tab is a tool that gets abandoned in a busy week.
Four: the scope was chosen because it was interesting. The chatbot on the website, the document classifier, the thing that demonstrates well. None of it touched the phone that rings out at two in the afternoon. Interest is a bad proxy for cost, and the tasks people find interesting to automate are almost never the tasks that are quietly losing the money.
Five: there was no way back. The old process was decommissioned on the day the new one went live, so the first bad week became an emergency instead of a decision. Firms that cannot fall back cannot experiment, and firms that cannot experiment end up defending a build nobody can prove rather than stopping it.
Why a small firm cannot run the kind of project those studies describe
The studies above describe a unit of work you cannot afford, and should not want. A six-month exploratory programme with a steering group is how a large company finds out whether something is worth doing. A firm of twenty people needs a smaller unit with a sharper edge, and needs it to produce an answer either way.
- One leak, named in advance. Not a department and not a category. The unanswered call, the next-day first reply, the slot that empties, the document somebody retypes. If two of them are bleeding, pick the larger and leave the other alone until this one is proven.
- One number, produced before the build starts. From your own call log, diary or inbox rather than from an industry average. It does not have to be exact. It has to be yours, and it has to be written down where both sides can see it, because it is the thing the result will be compared against.
- One acceptance test, written by the person paying. A sentence that makes this a success, in your words, before anyone builds anything. If the buyer cannot write that sentence, the project has no finish line and will be argued about instead of measured. The supplier writing it for you is the failure mode dressed as helpfulness.
- A way back, tested rather than promised. The old routing stays configured, somebody has actually switched back once during the build, and the fortnight it would cost to abandon the whole thing is a price you agreed to pay before you started.
That list is our delivery contract rather than a theory about other people's failures. It is what a first build looks like here: ten working days to a first working version, one leak, one number, an acceptance test in your words, and the old path still switched on underneath. We wrote it this way because these five modes are the ones we watched happen to firms before they called us.
The post-mortem, including the one where it did not work
A first project that cannot fail honestly is not an experiment, it is a purchase. So the end of one is a short written review with four questions, and it is written whether the result was good or bad. This is the part almost nobody does, and it is why so many firms have a drawer of automations nobody can account for.
- What was the number before, and what is it now? The same measure, taken the same way, over a comparable stretch of time. If the measurement changed halfway through, say so in the review rather than picking whichever version flatters the result.
- What did the machine hand back to a human, and why? Handovers are not failures, they are the design working. But the reasons are a map: if a quarter of calls hand over for the same reason, that reason is either the next rule to write or the edge of what should be automated at all.
- What broke, and how long until somebody noticed? The second half matters more than the first. A fault found in an hour is an incident. The same fault found in five weeks is a measurement problem, and the fix is monitoring rather than software.
- Expand, hold, or stop? All three are acceptable and only two of them are available to a firm that never wrote the number down. Stopping is not a wasted fortnight: it is the cheapest true answer you will get to a question that otherwise costs three years of not trusting any of it.
Which puts the whole argument back before the build, where it belongs. Most of what goes wrong is decided by the leak you pick and the number you write down, and both of those are readable from the outside. The €29 Check reads your call paths, reply paths, booking paths and document paths the same day and comes back with the candidates ranked and the arithmetic shown, so the first project is chosen on evidence rather than on which demonstration was the most convincing.