Insights · The 95% problem
The GenAI Divide, annotated: why 95% of AI pilots die — and the pattern in the 5% that live
TL;DR
- MIT Project NANDA's State of AI in Business 2025 found 95% of enterprise GenAI pilots deliver zero measurable P&L impact, against $30–40B of spend.
- Adoption isn't the problem: 80%+ of organizations piloted tools like ChatGPT or Copilot. Of those evaluating enterprise-grade systems, only ~5% reached production with value.
- External partnerships succeeded ~67% of the time — roughly twice the rate of internal builds.
- The report blames a "learning gap": systems that don't retain feedback or improve. We think that's downstream of something simpler — no dollar-valued target and no evaluation baseline. You can't learn against a scoreboard that doesn't exist.
- The 5% pattern: one workflow, back-office first, deep integration, a partner accountable to a number, and releases gated on measured improvement.
What is the "GenAI Divide" report?
In mid-2025, MIT Media Lab's Project NANDA published The GenAI Divide: State of AI in Business 2025 — the study behind the most-quoted number in enterprise AI. The research combined interviews with executives, surveys of employees, and analysis of roughly 300 public AI deployments. Its headline: despite $30–40 billion in enterprise GenAI spending, 95% of organizations were getting zero measurable business return.
The number traveled fast because it matched what operators were quietly seeing: pilots everywhere, impact nowhere. But most citations stop at the statistic. The report's internals — and its blind spots — are more useful than its headline.
What the report actually found
- The funnel is brutal at the enterprise tier. About 80% of organizations piloted consumer-grade assistants (ChatGPT, Copilot) and got diffuse individual productivity from them. But of firms evaluating enterprise-grade, task-specific systems, ~60% assessed one, ~20% piloted, and only ~5% reached production.
- Buying beats building — about 2:1. External partnerships and purchased tools succeeded ~67% of the time; internal builds succeeded roughly a third of the time.
- Budgets point at the wrong end of the business. Spending concentrates in sales and marketing, while the measured ROI shows up in unglamorous back-office work — operations, finance, document handling.
- A shadow AI economy already exists. Only ~40% of companies had official LLM subscriptions, yet ~90% of surveyed workers used personal AI tools daily. Employees crossed the divide before their employers did.
- Mid-market moves ~3× faster from pilot to implementation (~90 days) than large enterprises (9+ months).
MIT's diagnosis: the learning gap
Most GenAI systems do not retain feedback, adapt to context, or improve over time.
That's the report's core explanation, and it's real: people happily use a chat assistant for drafts, then reject the same model for mission-critical work because it forgets everything, repeats mistakes, and can't be trusted unattended. A static tool stays a demo forever.
Where the diagnosis stops short
Here's the annotation the report needs. "Systems that learn" is the visible difference between the 5% and the 95%. But learning is an effect, not a cause. A system can only improve against a definition of better — and that definition is exactly what failed pilots never write down.
In practice, the failed pilot goes like this: a team picks a technology ("we need agents"), builds a demo, shows it around, and only then asks how to prove it's working. There's no target metric, no baseline of current cost, no harness scoring outputs. Feedback has nowhere to accumulate — not because the model can't learn, but because nobody built the scoreboard it would learn against.
Read the buy-versus-build gap through the same lens. Vendors don't succeed 2:1 because their models are better — they largely use the same models. An external partner is contractually forced into measurement discipline: define success before the invoice, demo weekly against it, survive renewal conversations with evidence. Internal pilots can drift for quarters on enthusiasm. The 67% isn't a technology number; it's an accountability number.
Same for the mid-market speed advantage: 90 days to production isn't superior engineering — it's forced scope. One workflow, one owner, a number the CFO personally watches. Large enterprises fail slower because they can afford to.
The 5% playbook, made concrete
| What the report says | What it looks like in practice |
|---|---|
| Pick one process, integrate deeply | One workflow, priced in dollars before any code — what it costs today, what it's worth automated |
| Aim at the back office | Document handling, ops, finance — where "success" has an unambiguous unit cost |
| Systems that learn | An evaluation baseline on your data from day one; every prompt, model, and agent change gated on beating it |
| Partner over solo build | Whoever builds it, someone must be accountable to the number weekly — that's the vendor discipline worth importing even if you build in-house |
| Line managers over central labs | The person who owns the workflow owns the metric — not an innovation team three floors away |
How not to be the 95%: a five-question test
- Can you state the dollar value? "This workflow costs us $X/month today; automated it's worth $Y." If you can't write that sentence, stop.
- Does a baseline exist? A scored set of real cases, run against the current process, before the first demo.
- Is every release gated? New model, new prompt, new agent step — nothing ships unless the score moves.
- Can a non-engineer see the number? If the metric lives in a notebook, it doesn't exist. Dashboards, not vibes.
- Who is accountable weekly? A name, a demo, a number — the discipline that makes external partners succeed at twice the rate.
This test is the whole reason evaluation infrastructure — harnesses, regression gates, LLM-judged scoring — has quietly become the highest-leverage investment in applied AI. Not because evals are fashionable, but because they're the difference between a system that can learn and a demo that can't.
Every build should start with a dollar figure and end with a dashboard.
That's how we work — value scoped up front, an eval baseline before the build, releases gated on the score. If you're staring at a pilot that can't prove itself, we'll put a number on it.
Book a free 15-min consultSources
- MIT Project NANDA, The GenAI Divide: State of AI in Business 2025 (July 2025)
- Fortune, MIT report: 95% of generative AI pilots at companies are failing (Aug 2025)
- Forbes, MIT finds 95% of GenAI pilots fail because companies avoid friction (Aug 2025)
Figures are as reported by the study and its press coverage; methodology counts vary slightly between outlets. Where we editorialize, we say so.