MIT NANDA's State of AI in Business 2025 report found that 95% of enterprise generative AI pilots fail to deliver measurable profit-and-loss impact, while only about 5% reach production with real value. The cause is not weak AI models. Researchers call it a "learning gap": most pilots run generic tools that never integrate with the company's systems or adapt to its actual workflows, so they perform fine in a demo and produce nothing once they meet real operations. Two decisions separate the 5% from the 95%: buying from a specialized partner instead of building alone (roughly 67% success versus 22 to 33%), and starting in back-office automation instead of the flashier, harder-to-measure sales and marketing pilots that absorb most AI budgets.
If your pilot stalled or you are trying to avoid that outcome before you start, see how we run AI strategy and executive advisory engagements built around exactly this data. The full breakdown of why pilots fail, and what to do instead, is below.
What did the MIT study actually find?
MIT NANDA's "The GenAI Divide: State of AI in Business 2025" is the source of the now widely quoted 95% figure. The research is not a survey of opinions, it is grounded in 150 interviews with business leaders, a survey of 350 employees, and an analysis of 300 public AI deployments. The finding: about 5% of AI pilots achieve rapid, measurable revenue or cost impact. The other 95% stall, delivering little to no measurable effect on the P&L.
That does not mean 95% of companies saw nothing happen. Many pilots technically launched, technically ran, and technically produced output. What they did not do is move a real business number. A chatbot that answers questions nobody asked, or a drafting tool nobody trusts enough to skip the manual rewrite, is not a failure of ambition. It is a project that never closed the gap between "works in a demo" and "changes an outcome."
Why do AI pilots actually fail?
Because of a "learning gap," not a model problem. As Fortune's coverage of the report puts it, most enterprise tools fail "not because of the underlying models, but because they don't adapt, don't retain feedback and don't fit daily workflows." The point is worth sitting with: executives in the study typically blamed weak model performance when a pilot stalled. The underlying data told a different story. The models were rarely the bottleneck. The workflow around them was.
That distinction matters because it changes what you fix. A team that concludes "our AI wasn't good enough" goes shopping for a better model. A team that understands the real cause looks at whether the tool ever actually plugged into the systems people use every day, whether it kept learning from corrections instead of repeating the same mistake, and whether anyone redesigned the surrounding process or just dropped a tool into an unchanged one. As one line from the research puts it bluntly: "Technology doesn't fix misalignment. It amplifies it." An AI tool bolted onto a broken or disconnected process does not fix the process. It automates the dysfunction and calls it progress.
Picture the two versions side by side. Version one: a team licenses a general chatbot, points it at customer emails, and calls the pilot done. Nobody connects it to the order system, so it cannot check a real order status. Nobody reviews its answers for a month, so it keeps repeating the same wrong assumption about a shipping policy that changed. Three months in, agents quietly stop trusting it and go back to typing replies by hand. That is a 95% story. Version two: a team picks the same customer-email problem, but wires the tool into the order and billing systems on day one, reviews a sample of its answers weekly, and corrects the pattern the moment it drifts. Same starting technology, same use case, opposite outcome, because one team treated integration and correction as ongoing work and the other treated launch as the finish line.
Does buying beat building for AI projects?
Yes, by a wide and consistent margin. MIT found that purchasing an AI solution from a specialized vendor, or building it in partnership with one, succeeds roughly 67% of the time. Internal, DIY builds succeed at roughly one-third of that rate, with different cuts of the same underlying data putting internal success between about 22% and 33%. Either way, going it alone cuts your odds by two to three times.
The reason is not that internal teams are less capable. It is that a first-time internal build usually has no more experience closing the integration and workflow-adaptation gap than the tool itself does, while a specialized partner has already closed that exact gap on other projects, repeatedly, and knows where it tends to break. Company size compounds this. A separate cut of the MIT data found startups succeeded more often than large enterprises, in part because bigger companies carry entrenched processes that reflect existing bureaucracy and internal politics an AI tool cannot route around on its own. Mid-market companies reached implementation in about 90 days on average; large enterprises took closer to 9 months, often because more approval layers meant more chances for the pilot to drift from its original scope.
Buy-versus-build is also not really about ownership, it is about who is responsible for closing the learning gap over time. A purchased or partner-built tool that nobody keeps adapting fails just as reliably as an internal build. The advantage of the 67% path is that it usually comes bundled with someone whose job is to keep adapting it.
Why does back-office automation outperform sales and marketing AI?
Because more than half of AI budgets go to the wrong place. MIT found 50 to 70% of generative AI spend goes to sales and marketing tools, the pilots that are easiest to demo and tie to a board-level KPI. But the report's own finding is that the biggest, most measurable ROI sits in back-office automation: eliminating outsourced business process work, cutting external agency spend, and streamlining operations that were already running on repetitive, well-defined steps.
The pattern makes sense once you name it. Back-office work like invoice matching, claims review, or lead routing is high volume, repetitive, and has a clean, countable baseline: how many invoices, how many hours, how many dollars, before and after. Sales and marketing pilots are often the opposite: lower volume, harder to isolate from every other factor influencing a deal or a campaign, and easy to feel impressive in a demo while being nearly impossible to prove in a P&L. Named ranges from the research back this up directly: $2 to $10 million a year in BPO cost reduction from back-office automation, a 40% lift in lead qualification accuracy, and a 10% lift in customer retention from the use cases that did measure cleanly, all outcomes with a number attached, not a vibe.
Not sure where your highest-ROI use case actually is? See our breakdown on what to automate first, built on the same logic MIT's data supports: start where the volume is high, the process is repetitive, and the before-and-after is easy to measure.
What do the successful 5% do differently?
Seven things show up consistently across the projects that made it past the pilot stage and into measurable production value, per Forbes's synthesis of the MIT data:
| What they do | Why it works |
|---|---|
| Partner externally | External experts reach deployment roughly twice as often as internal-only teams. |
| Ground the initiative in a measurable strategy | A cross-division, numbers-first plan beats chasing whatever is trending. |
| Push adoption to the front line, keep accountability central | Front-line teams shape how the tool fits daily work; leadership keeps ownership of the outcome. |
| Integrate fully, not partially | Embedding into ERP, CRM, and finance systems beats a disconnected point solution every time. |
| Map the real bottleneck first | Most delays trace back to disorganized data or an inconsistent process, not a missing feature. |
| Manage the culture change on purpose | Training and workflow adaptation, not a tool drop and a memo. |
| Start in back office, not the spotlight | Operations, procurement, and finance show the clearest ROI, even if they demo less well. |
MIT's lead author, Aditya Challapally, summarized the pattern in the excelling minority in one line: they "pick one pain point, execute well, and partner smartly." Every one of the seven traits above is a specific version of that same discipline.
Want the buy side of that 67% without managing a vendor yourself? Hire AI agents already integrated into the workflow, with someone accountable for keeping them adapted as your business changes.
Has anything improved since the MIT report?
Not by mid-2026. A Forbes follow-up in April 2026 cites Gartner research finding that only 1 in 5 AI investments generate any measurable ROI, and only 1 in 50 deliver transformational value, a harsher read than MIT's original 5%, gathered roughly eight months later. PwC's CEO Survey from January 2026, covering 4,454 executives across 95 countries, found only 30% reported revenue increases from AI in the prior 12 months, and 56% said AI delivered zero cost or revenue improvement at all.
The same period produced real counterexamples worth naming, because they show the pattern still holds. Snowflake's AI-based sales training for 3,000 sellers saved more than 1,200 manager hours per quarter and an estimated $700,000 a year, a 4 to 5x return on the time and spend invested. JPMorgan saved close to $1.5 billion through AI-driven fraud prevention and operational efficiency. As Snowflake's Nathan Irby put it, "The clearest signal is almost always time returned to people. Hours saved, mapped to spend, gives you a ratio you can defend." Notice what both examples share with the MIT pattern: a specific, back-office-adjacent process, a clear before-and-after number, and a defensible ratio, not a company-wide AI initiative with no baseline.
Klarna's story is worth a mention precisely because it shows both sides of this at once. In 2024, Klarna reported its AI assistant automated roughly two-thirds of customer-service chats for a projected $40 million profit impact. By 2025 to 2026, Klarna also found the quality of that AI-handled work fell short of human staff and resumed hiring people for the interactions that needed it. Both facts are true. A pilot can show a real, measurable early win and still need ongoing adaptation to hold up at scale, which is exactly the "learning gap" MIT is describing: closing it once at launch is not the same as closing it continuously.
How to avoid becoming part of the 95%
A few concrete moves, drawn directly from what separates the two groups, apply before you spend a dollar on your next AI project.
- Pick one measurable pain point, not a platform. A named process with a before-and-after number beats a broad "AI transformation" initiative every time.
- Favor back office for the first win. High volume, repetitive, easy to measure. Save the harder-to-prove sales and marketing use cases for after you have a track record.
- Choose a partner over a solo internal build, unless you have a specific regulatory or IP reason not to. The success-rate gap is too large to ignore.
- Insist on integration, not a bolt-on. If the tool cannot connect to the systems your team actually works in every day, it will not survive contact with daily operations.
- Budget for ongoing adaptation, not just launch. The learning gap does not close once. It has to stay closed as your data, staff, and processes change.
- Measure hours or dollars, not activity. "We ran 10,000 conversations" is not a result. "We cut response time 40% and saved 1,200 hours a quarter" is.
How to get started
If your last pilot stalled, the MIT data suggests the reason is rarely the model you picked. It is more likely that the tool never integrated into how your team actually works, or that it launched in a flashy, hard-to-measure area instead of the high-volume back-office process where the ROI was always going to be clearest. Both are fixable, and neither requires starting over from scratch on a bigger model.
The fastest way onto the right side of the 95/5 split is the one the data itself points to: partner with a team that has already closed the integration and adaptation gap elsewhere, start in back office where the baseline is easy to measure, and keep adapting the tool after launch instead of treating go-live as the finish line. That is the whole model behind how we run projects. Book a free consultation below and we will find your highest-ROI use case and build the plan to land it in the 5%.
