All insights
The Market6 min read

The Pilot Graveyard

Vishal Sachar

Vishal Sachar

Co-Founder & CEO of CLRT

The most repeated statistic in enterprise AI is a misquote. The line that travelled says 95 percent of AI pilots fail. The study behind it, MIT Project NANDA's preliminary 2025 research, says something different and more damning: 95 percent of organisations are getting zero measurable return from generative AI. Pilots are not the subject of that sentence. Organisations are. The difference sounds pedantic and is anything but, because the two readings point at two different culprits. If pilots fail, the technology was not ready, and the fix is patience. If organisations get nothing, the question becomes what those organisations chose to point the technology at. Read the study that way and the graveyard stops looking like evidence against AI. It starts looking like evidence about how companies choose.

95%1
of organisations are getting zero measurable return from generative AI, per MIT Project NANDA's preliminary findings
MIT Project NANDA, 2025
5%2
of integrated AI pilots are extracting millions in value, the same study's rarely quoted other finding
MIT Project NANDA, 2025
40%+3
of agentic AI projects will be cancelled by the end of 2027, Gartner predicts
Gartner, 2025

Start with what the research actually found. MIT Project NANDA drew on 52 interviews, 153 survey responses, and more than 300 public AI initiatives, and was careful to label its own findings preliminary. Within that sample, about 60 percent of organisations evaluated enterprise AI tools, 20 percent got as far as a pilot, and 5 percent reached production. The same report notes that the small share of integrated pilots that do land are extracting millions in value. So the technology demonstrably works, for a few, at material scale. The funnel is the real finding: the attrition is not evenly spread bad luck, it is a shape, and shapes have causes. A number that collapses from 60 to 20 to 5 is not describing a tool that fails. It is describing a sequence of choices that were made long before any model was asked to do anything.

FIG. 01The attrition has a shape. Sixty percent of organisations evaluated enterprise AI tools, twenty percent reached a pilot, five percent reached production. MIT Project NANDA, preliminary findings, 2025.
01The selection effect

Look at which pilots get chosen and the shape explains itself. Companies pilot what is easy to pilot. Easy means visible: a chore everyone recognises and nobody defends. Easy means low stakes: if the output is wrong, nothing breaks and nobody is called. Easy means self-contained: no data dependencies, no compliance review, no other department's sign-off, nothing that requires a meeting before the demo. Those properties are precisely why the pilot can start on Monday. They are also precisely why it cannot move a P&L line, because the work that moves money is defended, consequential, and entangled with data and approvals by definition. The selection is not random and it is not stupid. It is what any organisation produces when piloting is delegated to whoever can begin without permission. The pilot did not fail in month three. It was pointed at nothing in month zero.

Plot every candidate pilot on two axes, ease of starting and value at stake, and the pattern becomes a map. The corner where pilots accumulate is easy and low value, and it fills itself, because it is the only corner that requires no judgment to enter. The corner where value lives is hard to start for reasons that have nothing to do with AI: the data sits in three systems, the workflow has an owner who must be convinced, the failure modes carry real cost. The market is now producing the evidence at scale. Meta's chief executive told staff in July 2026 that agents had not progressed as quickly as he had hoped, at a company expecting to spend roughly 145 billion dollars on AI infrastructure this year. Gartner predicts that more than 40 percent of agentic AI projects will be cancelled by the end of 2027. The graveyard is still filling.

FIG. 02Illustrative. Ease of starting and value at stake are independent axes, and most pilots are selected on only one of them. The graveyard corner is the only one that requires no judgment to enter.
02Choosing for value

A pilot chosen for value is built backwards. It starts from a line in the P&L that must move, or a cost that recurs, or a risk that is priced, and works upstream to the workflow that produces it. Before anything runs, it has an owner who wants the number to change, a baseline measured honestly, and a written definition of what failure looks like. It is harder to start, which is the point: the difficulty is not an obstacle on the way to the value, the difficulty is where the value has been hiding. Deloitte's 2026 survey of more than 3,000 organisations found 38 percent piloting and only 11 percent in production, a funnel with the same shape as NANDA's a year earlier, at larger scale. The organisations inside the 11 percent are not luckier or better resourced. They chose differently before they started.

FIG. 03Counted honestly: each square is an organisation, not a pilot. Five in one hundred report measurable return; the other 95 are unmoved P&L lines, not failed technology. MIT Project NANDA, preliminary findings, 2025.

That choice cannot be delegated to enthusiasm, and it cannot be bought from a vendor, because the vendor's demo is optimised for the same corner the graveyard occupies: visible, low stakes, self-contained. Choosing for value requires reading a specific business. Which lines actually move it, which workflows feed those lines, where the data really lives and how much of it can be trusted, what a wrong answer costs in each candidate process, and which owner will defend the change when it becomes inconvenient. None of that is technical knowledge, and all of it has to be settled before the technical work begins. The uncomfortable implication of the 95 percent is not that AI is overhyped. It is that most organisations spent a year proving they can deploy technology, and no time at all deciding where deploying it would matter.

The pilot did not fail in month three. It was pointed at nothing in month zero.

A deeper dive

The subtler cost of a graveyard pilot is not the money, which is usually small. It is the lesson it manufactures. An easy pilot that goes nowhere produces the conclusion that AI does not work here, and that conclusion is the most expensive possible output, because it is wrong, it is remembered, and it shapes the next budget cycle. The same NANDA research contains the evidence that the conclusion is false: only about 40 percent of the companies surveyed had bought an official LLM subscription, yet workers at more than 90 percent of them were regularly using personal AI tools for work. The organisation declares the technology unproven while its own staff adopt it privately at near-total penetration. That split should be read as a verdict on the pilots, not the tools. Employees point AI at their own friction, where they know exactly what the work is worth, and it sticks. The organisation points AI at whatever was easiest to approve, and it dies. Same models, same year, same company. The only variable is where each was pointed, and who did the pointing.

The selection discipline that separates the 5 percent is unglamorous and specific, which is why it does not spread by conference talk. It means starting from the P&L and walking backwards into the workflow, not starting from a capability and walking forwards into a use case. It means accepting that the best candidate usually crosses a departmental boundary, touches data that is messier than anyone admits, and needs an owner who has something to lose, all of which makes it slower to start and impossible to demo in week one. It means writing down the number the pilot must move, the baseline it is measured against, and the failure condition that kills it, before a single prompt is written, so that the pilot can actually teach something whichever way it goes. None of this is model work, and none of it improves with the next model release, which is exactly why the funnel has kept its shape while the technology underneath it improved dramatically. The scarce input is judgment about the business, applied before the engineering, and it is the input most pilot programmes budget at zero.

Work with CLRT

Deciding where to point AI before anything is built is the work CLRT does. We start from your P&L, not from a demo, and we will tell you honestly which workflows deserve a pilot, which are graveyard bait, and what the pilot must measure to prove anything at all. If you want that thinking applied to your own operation, CLRT Ascent runs the diagnostic at ascent.clrtstudio.com. And if you already have a graveyard of your own, bring us the pilot that mattered most. We will find where it should have been pointed.

Vishal Sachar

Vishal Sachar

Vishal Sachar is the Co-Founder and CEO of CLRT, where he helps UAE businesses make sense of applied agentic AI and put it to work. He writes on agentic systems, AI governance, and the economics of automation. Reach him at vishal@clrtstudio.com or on LinkedIn.

Start here

Skip the reading. See where your leverage leaks.

Ascent is our free diagnostic. Ten minutes, and you have the one workflow worth building first.