AI Operations · established evidence
Why AI Pilots Fail: Reading the MIT 95% Report for Small-Business Operators
The most-quoted number in business AI right now is that 95 percent of enterprise generative-AI pilots produced no measurable profit-and-loss impact. It comes from MIT's NANDA initiative, and the reason it matters is not the figure but the cause the researchers assigned to it. Asked why AI pilots fail, they did not find weak models. They found organizations that could not adopt the tools: systems that could not retain feedback, learn a specific business's context, or fit the way work already happened. That is a management finding, not a technology one, which is quietly good news for a small business. The report studied large enterprises, so the translation to a ten-person operation is an inference, not a proven transfer. But the failure modes it names are smaller, cheaper, and more visible at that scale than at ten thousand. This piece reads the report closely, separates what it establishes from what it only suggests, and turns the avoidable failures into things an owner can check before spending.
The number everyone quotes, and the number that matters
In August 2025, MIT's NANDA initiative published "The GenAI Divide: State of AI in Business 2025," built on more than 300 enterprise deployments, 52 structured case studies, and 153 leadership interviews. The headline that traveled was blunt: about 95 percent of the generative-AI pilots studied failed to produce a measurable profit-and-loss impact.
That figure did the rounds as proof that AI is overhyped. Read that way, it is almost useless. The load-bearing part of the report is not the 95 percent; it is the reason behind it. The researchers attributed the gap to organizational adoption failure rather than to model quality. The tools that stalled were not, in the main, too weak to do the work. They could not retain feedback, adapt to a specific business's context, or slot into an existing workflow, so they never crossed from demo to durable use.
That reframing changes who the report is for. If pilots failed because the underlying models were incapable, a small business could do nothing but wait for better technology. If they failed because of how the tools were chosen, wired in, and owned, then most of the failure is a set of decisions, and decisions are checkable before money is spent.
What the MIT report actually measured, and its limits
Using a study well starts with what kind of study it is. "The GenAI Divide" is a survey-and-case-study report, not a controlled experiment. It documents a strong pattern across a large sample of deployments and interviews, but it does not isolate cause the way a randomized trial would. Its conclusions are best treated as directional and well-evidenced, not as a proven causal law.
The second limit matters more for the reader of this article: the sample is enterprise. These are large organizations with budgets, procurement teams, and IT departments. A ten-person med-spa or a home-services crew is not a scaled-down version of that; it has different constraints and, often, a shorter distance between a decision and its consequences. Every time this piece carries a finding down to small-business scale, that move is a reasoned inference, and it is flagged as one.
Naming those limits is how you use a report instead of just repeating its headline. The failure modes below survive that scrutiny because they are structural, the kind of thing that gets worse, not better, when an organization is large and slow.
Adoption failure, not model failure: the mechanism
The report's central mechanism is what it calls a learning gap. The pilots that died tended to use tools that behaved the same on day ninety as on day one: they could not remember what a business corrected last week, could not absorb the context that makes an answer right for this business rather than a generic one, and could not bend to a process the team already ran. The model was fluent; the system around it did not learn.
This is why "ai adoption failure" is a more accurate label than "AI failure." The point of friction sat between a capable tool and an organization that could not metabolize it. The same report notes that a large share of the value that did land came from unglamorous places, back-office and process work where a tool could be pointed at a narrow, well-defined job, rather than from the ambitious, customer-facing launches that get announced.
For a small operator, the practical read is that the model on offer is rarely the variable that decides the outcome. The variables that decide it are whether the job is narrow enough to define, whether the data and tools underneath are clean enough to build on, and whether one person owns the process the tool is supposed to improve.
Buy versus build: the finding with the sharpest edge
One result in the report cuts cleanest for a small business. Purchased or partnered tools, systems bought from a specialist vendor, succeeded roughly 67 percent of the time, while internally built tools succeeded at roughly a third of that rate. Building your own was, in this sample, the losing move far more often than buying.
That runs against a common instinct that the "serious" path is to build something bespoke. For an organization without an engineering bench, the report's pattern suggests the opposite: a focused tool from someone who has already solved the boring, failure-prone parts beats a custom build that has to discover them.
Small businesses, as it happens, already lean the direction the evidence favors. A 2025 U.S. Chamber of Commerce survey of 3,870 small businesses found that 63 percent of AI adopters rely on external tools rather than building in-house, with only 8 percent fully in-house. The instinct is right; the report's caution is about the next step. Buying does not remove the adoption problem. A purchased tool bolted onto a messy process still fails. It shifts the hard work from building the tool to choosing the right one and wiring it into a process that can actually use it.
The adoption gap is measured across the whole field, not one report
A single influential study should never carry an argument alone. The useful thing about the MIT finding is that it sits inside a consistent picture drawn by independent researchers using different methods.
Stanford's 2026 AI Index found that while 88 percent of organizations now use AI in at least one business function, fewer than 10 percent have fully scaled it in any single function, and 74 percent now name inaccuracy as their top AI risk, up 14 points in a year. McKinsey's 2025 global survey, across nearly 2,000 respondents in 105 countries, found 23 percent reporting they had scaled an agentic AI system somewhere in the enterprise, yet in any given function no more than about 10 percent said they were scaling agents there. Adoption is broad and shallow at the same time.
The direction of travel is not in doubt. The same U.S. Chamber survey found small-business AI adoption rose from 23 percent in 2023 to 58 percent in 2025. More businesses are adopting; the gap the MIT report names is the distance between adopting and getting a result. That gap, not the technology, is the thing a small operator should be planning around.
What the deployments that worked had in common
It is easy to read a 95 percent failure figure as proof AI does not work. The best-documented workplace study of the technology says otherwise, and the contrast is instructive.
Brynjolfsson, Li and Raymond studied the staggered rollout of a generative-AI assistant to 5,179 customer-support agents at a software firm, published in the Quarterly Journal of Economics in 2025. Access raised productivity by 14 to 15 percent on average, with a 34 percent gain concentrated among novice and lower-skilled workers and near-zero effect on the most experienced. The mechanism was augmentation, not replacement: the tool spread the tacit best practices of the strongest workers to everyone else.
Two things carry across. First, the wins came from human plus AI on a narrow, high-volume task with a clear baseline to measure against, the same shape the MIT report saw succeed. Second, this study is enterprise too; no small-business replication of it exists, so reading a five-person front desk into a 5,000-agent contact center is an inference, not a proven transfer. The takeaway is directional: where AI has paid off in rigorous study, it did so by making people better at a defined job, not by being turned loose to run one.
The avoidable failure modes, checkable at small-business scale
Collapsing the evidence into practice, the failures that recur are not exotic. They are a short list of conditions an owner can inspect before committing a budget. None of these requires an IT department to see.
- No narrow job. The tool is aimed at a broad ambition ("handle the front desk") rather than one defined, repeatable task with a name. Broad scope is where the studied pilots died.
- A messy foundation. The data or the tools underneath are disorganized or do not connect, so the tool sits on sand. Published 2026 readiness frameworks agree this, not the model, is where most efforts stall.
- No owner. No single person is responsible for the process the tool is meant to improve, so corrections never get made and the system never learns.
- No baseline. There is no measured "before" number, which means there is no way to tell whether the tool worked, and no way to justify keeping it.
- Buying before diagnosing. A system is purchased to solve a problem that was never precisely defined, so the wrong tool arrives first and creates work instead of removing it.
- The autonomy leap. The tool is handed an irreversible action with no human check, rather than being kept to a reversible, supervised scope while trust is earned. Benchmark work on tool-using agents shows they still fail a large share of realistic tasks and, worse, behave inconsistently when the same task is repeated, which is exactly the case for a human check at the point where an action cannot be undone.
Why this is a diagnosis-before-purchase problem
Line the findings up and they point one direction. The technology is rarely the deciding variable. The deciding variables are the shape of the job, the state of the data and tools underneath, who owns the process, whether there is a baseline to measure against, and whether the first move is bought and narrow rather than built and sweeping.
Every one of those is knowable before a contract is signed, and cheaper to check than to fix afterward. That is the real lesson a small business should take from a 95 percent failure rate: the failures cluster in avoidable, pre-purchase decisions, and the order in which those decisions are made matters more than the enthusiasm behind them. A pilot that starts with a clear read of readiness is not guaranteed to succeed, but it removes the specific mistakes that sank most of the ones MIT studied.
The evidence
Key findings, with their sources
-
About 95% of enterprise generative-AI pilots studied produced no measurable profit-and-loss impact, and the gap was attributed to organizational adoption failure rather than model quality.
established MIT NANDA, "The GenAI Divide: State of AI in Business 2025" (300+ deployments, 52 case studies, 153 interviews), 2025.
-
Purchased or partnered AI tools succeeded roughly 67% of the time, while internally built tools succeeded at about a third of that rate.
established MIT NANDA, "The GenAI Divide: State of AI in Business 2025," 2025.
-
88% of organizations use AI in at least one business function, but fewer than 10% have fully scaled it in any single function; 74% now cite inaccuracy as their top AI risk, up 14 points in a year.
established Stanford HAI, "The 2026 AI Index Report," 2026.
-
23% of organizations report scaling an agentic AI system somewhere, yet in any single function no more than about 10% say they are scaling agents there.
established McKinsey & Company, "The State of AI: Global Survey 2025" (n=1,993 across 105 countries), 2025.
-
Small-business AI adoption rose from 23% (2023) to 58% (2025); 63% of adopters rely on external tools and only 8% build fully in-house.
established U.S. Chamber of Commerce, "Empowering Small Business" survey (3,870 U.S. small businesses), 2025.
-
A generative-AI assistant raised support-agent productivity by 14 to 15% on average, with a 34% gain for novices and near-zero for experts, by spreading the best workers' tacit practices, an augmentation effect, not full automation.
established Brynjolfsson, Li & Raymond, "Generative AI at Work," Quarterly Journal of Economics 140(2), 2025 (NBER WP 31161, 2023).
-
Even frontier tool-using agents succeed on fewer than half of realistic customer-service tasks, and repeating an identical task eight times, the same agent succeeds only about a quarter of the time.
established Yao et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction," arXiv:2406.12045, 2024 (Sierra Research).
Calibration
What is proven, what is promising, what is unproven
| Evidence tier | Tactics | What the evidence says |
|---|---|---|
| established | The 95% no-P&L-impact figure and the buy-versus-build result; the field-wide adoption gap (Stanford, McKinsey, U.S. Chamber); the peer-reviewed augmentation productivity finding. | MIT NANDA 2025; Stanford AI Index 2026; McKinsey State of AI 2025; U.S. Chamber 2025; Brynjolfsson/Li/Raymond, QJE 2025. |
| emerging | That the same avoidable failure modes are identifiable and checkable at ten-person scale, and that enterprise augmentation gains would transfer to a small front desk. | Reasoned inference from the enterprise studies above; no small-business replication exists, so treat as directional, not proven. |
| contested | Any precise dollar figure for what a failed pilot "costs" a specific small business. | Not established in the sources here; the studies report rates and mechanisms, not per-business loss amounts. Do not assert a number. |
Reference
Glossary
- AI pilot
- A limited trial of an AI tool on a real business problem, run to see whether it works before committing to it more broadly.
- P&L impact
- A measurable change in profit and loss, revenue earned or cost saved, that can be traced to the tool. The MIT report used this as the bar a pilot had to clear to count as a success.
- Adoption failure
- When a capable tool fails not because it is technically weak but because the organization cannot fit it into how work is done, retain its corrections, or give it the context to be useful.
- Augmentation versus automation
- Augmentation makes a person better at a defined task while they stay in the loop; automation removes the person. The best-evidenced workplace gains came from augmentation.
- Buy versus build
- The choice between purchasing an AI tool from a specialist vendor and building one in-house. In the MIT sample, buying succeeded far more often than building.
Straight answers
Frequently asked questions
What does the MIT report actually say about why AI pilots fail?
It found that about 95 percent of the enterprise generative-AI pilots it studied showed no measurable profit-and-loss impact, and that the cause was organizational adoption failure, not weak models. The tools that stalled could not retain feedback, learn a specific business's context, or fit an existing workflow. It is a survey-and-case-study report, so read its conclusions as strong and directional rather than as proven cause.
Does a 95% failure rate mean AI does not work for small businesses?
No. It means most pilots failed at adoption, not that the technology is useless. The best-evidenced workplace study of AI shows real productivity gains, largest for less-experienced workers, when a tool is aimed at a narrow, measurable job and a person stays in the loop. The failures cluster in avoidable decisions about scope, data, ownership and measurement, which is what makes them worth checking before you spend.
Is it better to buy AI tools or build them?
In the MIT sample, purchased or partnered tools succeeded roughly 67 percent of the time while internally built ones succeeded at about a third of that rate. For a business without an engineering team, buying a focused tool from a specialist tends to beat a custom build. Buying does not remove the work, though; a purchased tool wired into a messy process still fails, so the effort moves to choosing the right one and fitting it in.
How do I know if my business is ready for AI?
Readiness comes down to a few things you can inspect without an IT department: a narrow, named job for the tool to do; data and tools underneath that are clean enough to build on; one person who owns the process; a baseline number to measure against; and a first move that is bought and narrow rather than built and sweeping. A structured read across those dimensions before you buy is far cheaper than fixing a wrong purchase after.
What is the single most avoidable reason pilots fail?
Starting broad. The studied pilots that died tended to aim at sweeping ambitions rather than one defined, repeatable task. Pairing that with buying a system before the problem is precisely defined is the most common and most avoidable way to waste money, and both are decisions you can correct before any contract is signed.
Provenance
Sources
- MIT NANDA, "The GenAI Divide: State of AI in Business 2025," 2025 (established; survey and case-study report, treat conclusions as directional)
- Stanford HAI, "The 2026 AI Index Report," 2026 (established)
- McKinsey & Company, "The State of AI: Global Survey 2025," November 2025 (established)
- U.S. Chamber of Commerce, "Empowering Small Business" AI survey (3,870 U.S. small businesses), 2025 (established)
- Brynjolfsson, E., Li, D. & Raymond, L., "Generative AI at Work," Quarterly Journal of Economics 140(2), 2025; NBER Working Paper 31161, 2023 (established, peer-reviewed)
- Yao, S. et al., "tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains," arXiv:2406.12045, 2024, Sierra Research (established, open benchmark)arxiv.org
Every figure above is attributed to a real, dated source and tagged with its evidence tier. Where a claim could not be verified to a primary source, it is not stated as fact.