Skip to content

Why AI projects fail small firms, and what the ones that worked did differently

GUIDE · FOR THE SCEPTIC · 11 MIN

Facts on this page verified August 2026.

The short answer: almost never the model, almost always the setup

They fail on decisions made before any software is chosen. Four causes account for most of it. Nobody wrote down the number the project was supposed to move. Nobody owned it once it was live. The machine sat beside the real work instead of inside it. And the scope was picked because it was interesting rather than because it was bleeding. All four are organisational, all four are cheap to fix at twenty people, and none of them is a question about which model you use.

The number everyone quotes, read properly

You have seen the headline: ninety-five percent of corporate AI projects return nothing. It comes from one study, it is a real finding, and almost every article repeating it gets the scope wrong. Here is what it measured.

95% · 5%

Corporate generative-AI projects with no measurable profit-and-loss impact, against the share that produced one.study, enterprise scaleSource: MIT Project NANDA, via Fortune, 2025

What it counted was large organisations, not firms like yours. One hundred and fifty interviews with leaders, a survey of three hundred and fifty employees, and three hundred publicly disclosed deployments, all at corporate scale. It is a finding about generative-AI programmes inside big companies. It is not a finding that the technology does not work, and it is not a prediction that a twenty-person firm putting an agent on its phone will fail nineteen times out of twenty.

The second number is the one people attribute to the wrong publisher. Search for AI failure rates and you will be told that RAND found more than eighty percent of AI projects fail. Open the report and the sentence is in the introduction, with a footnote to a magazine commentary.

>80%

AI projects said to fail, about twice the rate of technology projects that do not involve AI.estimate RAND quotes, not RAND researchSource: RAND Corporation, 2024

What RAND actually did is more useful than the number it borrowed. Sixty-five interviews with data scientists and engineers of at least five years standing, conducted between August and December 2023, producing five root causes. Four of the five are organisational: the wrong problem, the wrong data, chasing the technology instead of the user, and infrastructure that cannot carry a deployed model. Only the fifth is about the limits of the technology itself.

The third number says the failures are being noticed, and acted on. Abandonment climbed sharply in one year, which is a healthier signal than it sounds: firms are now stopping things that are not working instead of leaving them switched on.

42% · 17%

Companies that abandoned most of their AI initiatives in 2025, against the same measure a year earlier.survey, 1,000+ respondentsSource: S&P Global Market Intelligence, via CIO Dive, 2025

Read the three together and the shape is clear. Failure is common, it is rising, and it is overwhelmingly organisational. Every one of these studies looked at companies far larger than yours, which is a limit on what they prove and also the opening: the causes they name are cheaper to avoid at twenty people than at twenty thousand, because at twenty people the person who decides is the person who watches it run.

The five ways it actually goes wrong

Each mode below is what it looks like from inside the firm, followed by what the projects that worked did instead. None of them is exotic. Every one of them is decided in the first fortnight, usually in a conversation nobody wrote down.

One: there was never a before-number. The project was approved on a feeling that something was slow, and it is judged a year later on a different feeling. Nobody can say what the phone was costing before, so nobody can say what it costs now, and the line item quietly loses the next budget argument to something that can show a figure. This is the most common of the five and the cheapest to have avoided.

Two: nobody owned it after the launch. The build finished, the supplier left, and the automation became everybody's and therefore nobody's. Six weeks later an opening hour changed, or a price list moved, and the agent started giving an answer that used to be right. Nobody had the job of noticing, so the fix was to switch it off.

Three: it sat beside the work instead of inside it. The machine produced something correct and then somebody had to copy it into the system where the work actually happens. That copying is the tax that kills adoption: it is small enough that nobody complains and large enough that people stop bothering within a month. A tool that lives in a separate tab is a tool that gets abandoned in a busy week.

Four: the scope was chosen because it was interesting. The chatbot on the website, the document classifier, the thing that demonstrates well. None of it touched the phone that rings out at two in the afternoon. Interest is a bad proxy for cost, and the tasks people find interesting to automate are almost never the tasks that are quietly losing the money.

Five: there was no way back. The old process was decommissioned on the day the new one went live, so the first bad week became an emergency instead of a decision. Firms that cannot fall back cannot experiment, and firms that cannot experiment end up defending a build nobody can prove rather than stopping it.

Why a small firm cannot run the kind of project those studies describe

The studies above describe a unit of work you cannot afford, and should not want. A six-month exploratory programme with a steering group is how a large company finds out whether something is worth doing. A firm of twenty people needs a smaller unit with a sharper edge, and needs it to produce an answer either way.

  1. One leak, named in advance. Not a department and not a category. The unanswered call, the next-day first reply, the slot that empties, the document somebody retypes. If two of them are bleeding, pick the larger and leave the other alone until this one is proven.
  2. One number, produced before the build starts. From your own call log, diary or inbox rather than from an industry average. It does not have to be exact. It has to be yours, and it has to be written down where both sides can see it, because it is the thing the result will be compared against.
  3. One acceptance test, written by the person paying. A sentence that makes this a success, in your words, before anyone builds anything. If the buyer cannot write that sentence, the project has no finish line and will be argued about instead of measured. The supplier writing it for you is the failure mode dressed as helpfulness.
  4. A way back, tested rather than promised. The old routing stays configured, somebody has actually switched back once during the build, and the fortnight it would cost to abandon the whole thing is a price you agreed to pay before you started.

That list is our delivery contract rather than a theory about other people's failures. It is what a first build looks like here: ten working days to a first working version, one leak, one number, an acceptance test in your words, and the old path still switched on underneath. We wrote it this way because these five modes are the ones we watched happen to firms before they called us.

The post-mortem, including the one where it did not work

A first project that cannot fail honestly is not an experiment, it is a purchase. So the end of one is a short written review with four questions, and it is written whether the result was good or bad. This is the part almost nobody does, and it is why so many firms have a drawer of automations nobody can account for.

  • What was the number before, and what is it now? The same measure, taken the same way, over a comparable stretch of time. If the measurement changed halfway through, say so in the review rather than picking whichever version flatters the result.
  • What did the machine hand back to a human, and why? Handovers are not failures, they are the design working. But the reasons are a map: if a quarter of calls hand over for the same reason, that reason is either the next rule to write or the edge of what should be automated at all.
  • What broke, and how long until somebody noticed? The second half matters more than the first. A fault found in an hour is an incident. The same fault found in five weeks is a measurement problem, and the fix is monitoring rather than software.
  • Expand, hold, or stop? All three are acceptable and only two of them are available to a firm that never wrote the number down. Stopping is not a wasted fortnight: it is the cheapest true answer you will get to a question that otherwise costs three years of not trusting any of it.

Which puts the whole argument back before the build, where it belongs. Most of what goes wrong is decided by the leak you pick and the number you write down, and both of those are readable from the outside. The €29 Check reads your call paths, reply paths, booking paths and document paths the same day and comes back with the candidates ranked and the arithmetic shown, so the first project is chosen on evidence rather than on which demonstration was the most convincing.

Every source on this page

Each claim above is numbered to one of these. Open them and check.

  1. 1. MIT Project NANDA, via Fortune: The GenAI Divide: State of AI in Business 2025 (2025)

    research

    150 interviews with leaders, a survey of 350 employees and 300 publicly disclosed deployments, all at corporate scale. It measures profit-and-loss impact of generative-AI programmes inside large companies. It says nothing about owner-run firms. The original PDF address on the MIT project domain now redirects rather than serving the file, so this reporting is the record we can point you at.

  2. 2. RAND Corporation: The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (Ryseff, De Bruhl and Newberry) (2024)

    research

    RAND interviewed 65 data scientists and engineers with at least five years of experience, between August and December 2023, and identified five root causes. The widely quoted "more than 80 percent of AI projects fail" appears in RAND's introduction with a footnote to a magazine commentary: it is a figure RAND repeats, not a figure RAND measured. The five causes are the part to read.

  3. 3. S&P Global Market Intelligence, via CIO Dive: AI project failure rates are on the rise (2025)

    research

    A survey of more than 1,000 respondents in North America and Europe, reported in March 2025. The same survey found the average organisation scrapped 46% of its proofs of concept before they reached production. Corporate respondents again, so the level is theirs and only the direction is likely to be yours.

We tried a chatbot and it went badly. Is this different?

Usually yes, and for a boring reason. A website chatbot answers questions that were not blocking a sale, sits outside the systems the work happens in, and has no number attached to it, which is three of the five failure modes above in one product. The question to ask about anything you are offered next is what it can finish on its own and which number it is supposed to move.

How do you measure whether it worked?

Against the number you wrote down before it was built, using the same method. For a phone build that is calls offered, calls answered and bookings made, compared over comparable weeks. We also count handovers and the reasons for them, because a machine that quietly guesses in order to look competent will show a good answer rate and a bad booking rate.

What is the fallback if it breaks?

The path you had before, which stays configured rather than being decommissioned on go-live day. Switching back is a change to the routing and takes about a minute, and somebody at your firm does it once during the build so it is a rehearsed step rather than a promise. You keep the transcripts either way.

Who owns the workflow afterwards?

You do, and that is a contractual answer rather than a friendly one: the accounts, the number, the data and the rules are yours. Inside your firm one named person owns it day to day, reads the transcripts weekly for the first month and can change the rules. Automation that only its supplier can edit is a dependency, not an asset.

How small can a first project be?

Small enough that abandoning it costs you a fortnight rather than a quarter. One trigger, one action, one measure: when a call goes unanswered, send a text with a booking link, and count the bookings. If a supplier tells you the smallest useful version is a six-month programme, they are describing their own commercial needs rather than yours.

Read next

What should a small firm automate first?

The selection test that stops mode four, run before anybody quotes you for anything.

Where every number on this site comes from

The register: which figures are research, which are ours, and what each one does not prove.

Missed-call cost calculator

The fastest way to produce the before-number that mode one is missing.

The €29 Check

Your leaks read and ranked the same day, with the arithmetic shown, before anything is built.

Pick the first project on evidence, not on the demonstration.

The Check reads your call paths, reply paths, booking paths and document paths, then ranks the candidates for your firm with the arithmetic shown.

Get the €29 Check

€29 once · same day, usually within a few hours · not a subscription