Insights · The 95% problem

The GenAI Divide, annotated: why 95% of AI pilots die — and the pattern in the 5% that live

Atlas Research · July 2026 · 9 min read · Sources at the end

TL;DR

  • MIT Project NANDA's State of AI in Business 2025 found 95% of enterprise GenAI pilots deliver zero measurable P&L impact, against $30–40B of spend.
  • Adoption isn't the problem: 80%+ of organizations piloted tools like ChatGPT or Copilot. Of those evaluating enterprise-grade systems, only ~5% reached production with value.
  • External partnerships succeeded ~67% of the time — roughly twice the rate of internal builds.
  • The report blames a "learning gap": systems that don't retain feedback or improve. We think that's downstream of something simpler — no dollar-valued target and no evaluation baseline. You can't learn against a scoreboard that doesn't exist.
  • The 5% pattern: one workflow, back-office first, deep integration, a partner accountable to a number, and releases gated on measured improvement.

What is the "GenAI Divide" report?

In mid-2025, MIT Media Lab's Project NANDA published The GenAI Divide: State of AI in Business 2025 — the study behind the most-quoted number in enterprise AI. The research combined interviews with executives, surveys of employees, and analysis of roughly 300 public AI deployments. Its headline: despite $30–40 billion in enterprise GenAI spending, 95% of organizations were getting zero measurable business return.

The number traveled fast because it matched what operators were quietly seeing: pilots everywhere, impact nowhere. But most citations stop at the statistic. The report's internals — and its blind spots — are more useful than its headline.

What the report actually found

  • The funnel is brutal at the enterprise tier. About 80% of organizations piloted consumer-grade assistants (ChatGPT, Copilot) and got diffuse individual productivity from them. But of firms evaluating enterprise-grade, task-specific systems, ~60% assessed one, ~20% piloted, and only ~5% reached production.
  • Buying beats building — about 2:1. External partnerships and purchased tools succeeded ~67% of the time; internal builds succeeded roughly a third of the time.
  • Budgets point at the wrong end of the business. Spending concentrates in sales and marketing, while the measured ROI shows up in unglamorous back-office work — operations, finance, document handling.
  • A shadow AI economy already exists. Only ~40% of companies had official LLM subscriptions, yet ~90% of surveyed workers used personal AI tools daily. Employees crossed the divide before their employers did.
  • Mid-market moves ~3× faster from pilot to implementation (~90 days) than large enterprises (9+ months).

MIT's diagnosis: the learning gap

Most GenAI systems do not retain feedback, adapt to context, or improve over time.

That's the report's core explanation, and it's real: people happily use a chat assistant for drafts, then reject the same model for mission-critical work because it forgets everything, repeats mistakes, and can't be trusted unattended. A static tool stays a demo forever.

Where the diagnosis stops short

Here's the annotation the report needs. "Systems that learn" is the visible difference between the 5% and the 95%. But learning is an effect, not a cause. A system can only improve against a definition of better — and that definition is exactly what failed pilots never write down.

In practice, the failed pilot goes like this: a team picks a technology ("we need agents"), builds a demo, shows it around, and only then asks how to prove it's working. There's no target metric, no baseline of current cost, no harness scoring outputs. Feedback has nowhere to accumulate — not because the model can't learn, but because nobody built the scoreboard it would learn against.

Read the buy-versus-build gap through the same lens. Vendors don't succeed 2:1 because their models are better — they largely use the same models. An external partner is contractually forced into measurement discipline: define success before the invoice, demo weekly against it, survive renewal conversations with evidence. Internal pilots can drift for quarters on enthusiasm. The 67% isn't a technology number; it's an accountability number.

Same for the mid-market speed advantage: 90 days to production isn't superior engineering — it's forced scope. One workflow, one owner, a number the CFO personally watches. Large enterprises fail slower because they can afford to.

The 5% playbook, made concrete

What the report saysWhat it looks like in practice
Pick one process, integrate deeplyOne workflow, priced in dollars before any code — what it costs today, what it's worth automated
Aim at the back officeDocument handling, ops, finance — where "success" has an unambiguous unit cost
Systems that learnAn evaluation baseline on your data from day one; every prompt, model, and agent change gated on beating it
Partner over solo buildWhoever builds it, someone must be accountable to the number weekly — that's the vendor discipline worth importing even if you build in-house
Line managers over central labsThe person who owns the workflow owns the metric — not an innovation team three floors away

How not to be the 95%: a five-question test

  1. Can you state the dollar value? "This workflow costs us $X/month today; automated it's worth $Y." If you can't write that sentence, stop.
  2. Does a baseline exist? A scored set of real cases, run against the current process, before the first demo.
  3. Is every release gated? New model, new prompt, new agent step — nothing ships unless the score moves.
  4. Can a non-engineer see the number? If the metric lives in a notebook, it doesn't exist. Dashboards, not vibes.
  5. Who is accountable weekly? A name, a demo, a number — the discipline that makes external partners succeed at twice the rate.

This test is the whole reason evaluation infrastructure — harnesses, regression gates, LLM-judged scoring — has quietly become the highest-leverage investment in applied AI. Not because evals are fashionable, but because they're the difference between a system that can learn and a demo that can't.

Every build should start with a dollar figure and end with a dashboard.

That's how we work — value scoped up front, an eval baseline before the build, releases gated on the score. If you're staring at a pilot that can't prove itself, we'll put a number on it.

Book a free 15-min consult

Sources

Figures are as reported by the study and its press coverage; methodology counts vary slightly between outlets. Where we editorialize, we say so.