Economics · 26 June 2026

Why most AI pilots fail, and what the survivors do differently

The technology usually works. The deployment usually does not. The distinction decides which side of the 95 per cent you land on.

In 2025, researchers at MIT's Media Lab looked at hundreds of enterprise generative AI deployments and reached a conclusion that has been quoted in boardrooms ever since: roughly 95 per cent of pilots produced no measurable impact on the profit and loss account. Not a modest impact. None that could be found.

The reflex is to blame the models. That is almost never the story. The models in those failed pilots were the same models running in the successful five per cent. What differed was everything around them.

How pilots die

They die in familiar ways. The pilot is built for the demo rather than the workflow, so it performs beautifully on the ten examples in the deck and collapses on the two thousand real cases with their missing fields and angry footnotes. Nobody wires it into the systems where the actual work lives, so a person spends their day copying results between windows, which is the job the pilot was meant to remove. Nobody owns it, so when it misbehaves in week three there is no one whose problem it is. There is no governance, so nobody is willing to let it touch anything that matters, which means it never does anything that matters. And nobody defined success at the start, so the pilot cannot fail, which is another way of saying it cannot succeed.

Any one of these is survivable. Most pilots have all five.

What the five per cent do

The survivors are unglamorous. They pick one process, not a platform ambition. They measure the current cost and error rate of that process before building anything, because a baseline you did not record is a baseline you will argue about later. They connect to the real systems in the first weeks, not as a phase-two promise. They ship something into shadow production early, running on live inputs while a person still does the job, so the comparison is data rather than opinion.

They keep humans at the gates that matter and let the system run free where the stakes are low, which is what makes the organisation willing to trust it with real work. And they treat exceptions as the curriculum: every case the system gets wrong becomes a rule, a check, or a better prompt, so the error rate falls month on month instead of being discovered all at once in an audit.

The industry has noticed

It is telling that the AI labs themselves have moved this way. The enterprise ventures announced in 2026 are built around engineers embedded in the customer's operation until the system runs, a model Palantir spent two decades proving. The lesson from both the research and the money is the same one. The model is not the product. The working deployment is.

If you are planning a pilot, the checklist is short. One process. A recorded baseline. Real integrations. Shadow mode before production. Gates where the risk is. An owner with authority. Do those six things and you are no longer playing the odds the MIT study describes.

Find out in thirty minutes

A free call with one purpose: a straight answer on whether agentic workflows fit your business. If they do not, we will say so.

Book a 30-minute fit call