One number has been pinned to every AI steering committee agenda this year. MIT's Project NANDA found that 95% of generative AI pilots at large enterprises produced no measurable impact on the bottom line.
It is a real finding, and it deserves the attention it got. But most people are reading it wrong. They hear it as a verdict on the technology. It is closer to a verdict on the process that surrounds the technology.
Look at what the number actually measures. It measures whether a company could demonstrate a financial result. That is a different question from whether one occurred.
Critics note the 95% figure counts only direct P&L impact within a short window, and that the sample skews toward bespoke, workflow-embedded projects. That's exactly the point: the narrow definition of "success" is itself part of the problem.
Three things go wrong, and none of them are the model
Nobody wrote down what success would look like.
Most pilots launch without predefined success criteria. That single fact makes the outcome inevitable. If you never defined the target, you cannot declare a hit, even in the case where the technology performed exactly as designed. The pilot ends, everyone has a positive impression and no evidence, and the finance team correctly refuses to fund the next phase.
The metrics collected were the easy ones.
Early enterprise AI reporting was built on usage: seats assigned, hours logged, teams onboarded. Those numbers are simple to gather and satisfying to present. They are also unrelated to the only question that matters, which is whether the work came out better than what it replaced.
The unglamorous 80% was never scoped.
Moving from a working pilot to production is mostly data engineering, governance, workflow integration, and measurement infrastructure. The model portion is small. IDC found that 88% of AI pilots never reach production at all, and that the failures cluster around governance, data readiness, and observability rather than model quality.
That last point deserves emphasis. The failures are not happening where people are looking.
The plateau nobody expected
If this were a temporary adjustment, we would expect it to be improving. It is not.
The fifth annual Domino Enterprise AI Report, published in July, surveyed 639 senior enterprise AI leaders. The share of enterprises whose returns fail to outpace their investment held at 57%, unchanged from 2025. Over the same period, 93% reported improved production capability, up from 88%.
Read those two lines together. Capability went up. Returns did not move. That is a two-year plateau, and it is the clearest signal in the data that the constraint is not technical.
S&P Global found that 42% of companies abandoned most of their AI projects in 2025, more than double the year before. Gartner expects more than 40% of agentic AI projects to be cancelled by 2027, citing costs, unclear returns, and weak risk controls.
What to do instead of another pilot
We would go further than the usual advice. The problem is the pilot as a format.
A pilot is designed to answer "can this work?", a question that stopped being interesting in about 2024. The answer is yes, in a demo, almost always. Then the demo ends and nothing changes.
The alternative is smaller and more boring, and it works:
Pick one workflow, not one department. A workflow has inputs, outputs, and a person accountable for the result. A department has politics.
Write the baseline down before you start. Cycle time, error rate, cost per unit, whatever fits. If you cannot measure it today, you will not be able to prove you improved it tomorrow.
Run it for a full quarter. One workflow-level metric, against a defined baseline, over at least a quarter, gives leadership something defensible. Two weeks of enthusiasm gives them nothing.
Budget for the integration, not the license. If 80% of the work is plumbing, a plan that funds only the model is a plan to stall.
Decide in advance what would make you stop. Most organizations have no mechanism to kill an AI project. So the projects do not die, they linger, consume budget, and quietly become the reason the next proposal is refused.
The part people do not like
Some of that 95% deserved to fail. Not every workflow should be automated. Some tasks are rare, or high-stakes, or so entangled with judgment that a faster version is a worse version.
An honest measurement program will tell you that. That is a feature. A company that can prove three things did not work has learned more than a company with thirty pilots and one anecdote.
The uncomfortable conclusion is that the AI question and the operating question have collapsed into the same question. You are not deciding whether the technology is good enough. You are deciding whether your organization can tell the difference between activity and results.
Most cannot yet. That is the actual gap, and it is a solvable one.
Research notes
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025): 95% of the 300+ large-enterprise initiatives reviewed showed no measurable P&L impact; the report points to measurement and workflow-integration failures rather than model capability.
- IDC research with Lenovo (Feb–Mar 2025): 88% of AI pilots did not reach production; for every 33 proofs of concept, four reached production, with governance, data readiness, and observability as recurring barriers.
- Domino’s fifth annual Enterprise AI Report (July 2026): 57% of 639 senior enterprise AI leaders reported returns that did not outpace investment, unchanged from 2025; 93% reported improved production capability, up from 88%.
- S&P Global Market Intelligence, Voice of the Enterprise: AI and ML Use Cases 2025: 42% of companies abandoned most AI projects in 2025, up from 17% in 2024; 46% of proofs of concept were scrapped on average before deployment.
- Gartner (June 2025) forecast that more than 40% of agentic AI projects will be cancelled by 2027, citing costs, unclear business value, and inadequate risk controls.
- Other surveys report higher success rates: McKinsey reported 34% measurable business impact, Deloitte 42% meeting or exceeding predefined criteria, Gartner 38% measurable ROI within the first year, and Metrigy reported that 85% of companies had fewer than 40% of pilots fail to continue to production.
- Common methodology considerations include direct-financial-ROI definitions, self-reported outcomes, non-random large-enterprise samples, short observation windows, and differing definitions of measurable impact.
- Documented deployment outcomes include Salesforce’s $100M annualized cost savings, IBM Client Zero’s $4.5B annualized productivity savings, ABB’s 16,200+ automated work items, CAIXA’s 80% productivity gains, and field experiments reporting up to 16.3% sales increases in online retail.