← Back to research
Briefing · Adoption· Updated 14 September 2026· 7 min read

The 95% problem: why most AI pilots never reach the P&L

An MIT report said 95% of company AI pilots produced no measurable return. Gartner put generative AI in the trough of disillusionment the same month. What the studies measured, what they missed, and why the 5% look the way they do.

A row of six small terracotta pots on a pale desk, one with a seedling

In August 2025 a single number did the rounds of every board in the country: 95% of generative AI pilots fail. It came from a report by MIT's NANDA initiative, it moved AI stocks on the day it was reported, and it was quoted far more often than it was read.[1][2] This briefing sets out what the report actually measured, what the other surveys of the same period say, and what the minority of deployments that do reach the profit-and-loss account have in common. It is written at the start of September with the sources available then; an update at the end carries what the following year's surveys found.

What the MIT report measured

The report, titled The GenAI Divide, drew on 52 structured interviews with enterprise stakeholders, a survey of 153 senior leaders and an analysis of 300 publicly announced AI initiatives, gathered over the first half of 2025.[1] Its headline finding was that despite $30 to $40bn of enterprise spending on generative AI, 95% of organisations were getting no measurable return, defined as no discernible impact on the profit-and-loss account within about six months of a pilot. The remaining 5% had extracted millions in value.[1][2] Early press accounts gave the sample as 150 interviews and 350 employees; the figures above are the report's own, and the discrepancy became one of the first things critics noticed.[2]

The more useful content is the explanation. The report argues that the tools which stall are the ones that do not learn: they do not retain feedback, adapt to context or improve over time, so a general-purpose chatbot that impresses in a demo is abandoned once it fails on a real workflow it cannot remember. Two patterns follow from that. Externally purchased tools and partnerships reached deployment about 67% of the time against about 33% for internal builds, largely because vendors had already done the integration work. And budgets were pointed at the wrong end of the business: more than half of generative AI spend went to sales and marketing, while the clearest returns came from back-office automation, where the work is repetitive and the savings are measurable.[1]

95%
Organisations with no measurable P&L impact from generative AI pilots, per MIT NANDA [1]
67% vs 33%
Deployment rate for purchased tools against internally built ones [1]
90% vs 40%
Workers using personal AI tools for work, against employers that had bought any [1]

What it did not measure

Three cautions. The first is about the definition. "No measurable P&L impact within six months" is a high bar that most software projects would also fail; it says as much about companies' failure to set a baseline before a pilot as about the pilot itself. The second is about the sample. Three hundred initiatives that had been announced publicly, and 153 leaders willing to answer a survey, is not a random draw of businesses, and the direction of the bias is unknown. The third is about the headline. The report did not find that AI does not work; it found that most companies had not yet organised themselves to capture the value, and its own 5% were doing very well.[1] Read that way, the number is a management finding rather than a technology one, which is more useful and less dramatic.

The trough, on schedule

Two weeks before the MIT report, Gartner's annual Hype Cycle for artificial intelligence placed generative AI in the trough of disillusionment, the stage where early demonstrations have given way to the difficulty of scaling into production, and put AI agents and "AI-ready data" at the peak of inflated expectations.[3] The sequencing is worth noticing. The technology that boards were disappointed by in August 2025 was the one they had been most excited about in 2023, and the one they were most excited about in August 2025 was the one that had not yet been through the same cycle. Sam Altman, whose company had the most to gain from the excitement, said in the same fortnight that he thought investors as a whole were overexcited.[4]

The adoption numbers say something different

Adoption was not the problem. McKinsey's global survey published in March 2025 found 78% of organisations using AI in at least one business function, up from 55% a year earlier, and 71% regularly using generative AI.[5] Eurostat's official statistics for 2024, which count enterprises of ten or more employees using any AI technology, put the EU average at 13.5%, up from 8.0% in 2023, with Denmark highest at 27.6% and Ireland at 14.9%, a little above the average.[6] The two figures measure different things, self-reported use by mostly large companies against a statistical census of all firms above ten staff, but they point the same way: a large majority of big companies and a growing minority of all companies were using AI somewhere. The MIT finding sits on top of that. Use was wide, and shallow.

The shallowness shows in the shadow numbers. In the MIT sample, 90% of employees used personal AI tools for work while only 40% of their companies had purchased a subscription.[1] Ibec's survey of the Irish workforce, published as this briefing was written, found 40% of employees using AI in their work, double the 19% of a year earlier, with 81% saying they could use it better with training and 27% saying they had received none.[7] The adoption was happening from the bottom, on personal accounts, outside any policy, which is both the reason the pilots underperformed and a data-protection problem in its own right.

Two cautionary tales

Klarna is the clearest case of the year. In February 2024 the payments company said its AI assistant was handling two-thirds of customer-service chats and doing the work of 700 agents.[8] In May 2025 its chief executive, Sebastian Siemiatkowski, told Bloomberg the company had cut too far: "Cost unfortunately seems to have been a too predominant evaluation factor," he said, "what you end up having is lower quality." Klarna began hiring human agents again on a flexible model, while the assistant kept handling the majority of routine conversations.[9][10] The lesson is not that the AI failed; it is that the metric was wrong, and the business found out from its customers.

The second is about who gets hired. A Stanford paper published on 26 August, using payroll records from ADP covering millions of American workers, found that employment of 22 to 25 year olds in the occupations most exposed to AI had fallen 13% relative to other groups since late 2022, while older workers in the same occupations and young workers in less exposed ones were unaffected. The effect ran through reduced hiring rather than redundancies, and was concentrated where AI substitutes for a task rather than assists with it.[11] It is one country and one dataset, and the authors were careful about causation, but it was the first hard evidence that the pilots were changing something even where they were not reaching the P&L.

What the 5% do

  • They start with one workflow that has a measurable cost, usually in the back office, rather than a company-wide tool rollout.[1]
  • They buy or partner where a vendor has already solved the integration, and build only what is specific to them.[1]
  • They set the baseline before the pilot, so that the return is measurable at all.[1]
  • Adoption is led by the line manager who owns the process, not by a central innovation team with no stake in the outcome.[1]
  • They give staff an approved tool and a policy, which turns the shadow usage into something the company can see and improve.[1][7]

Some companies were already running the opposite experiment from the top. Shopify's chief executive told staff in April that using AI was now a baseline expectation and that teams would have to show why AI could not do a job before asking for headcount.[12] Whether a mandate produces the 5% pattern or just more pilots is the question the next year of data would answer.

Sources

  1. [1]The GenAI Divide: State of AI in Business 2025 · MIT NANDA · Jul 2025
  2. [2]MIT report: 95% of generative AI pilots at companies are failing · Fortune · 18 Aug 2025
  3. [3]Gartner Hype Cycle identifies top AI innovations in 2025 · Gartner · 5 Aug 2025
  4. [4]Sam Altman's paradox: warning of an AI bubble while raising trillions · Fortune · 19 Aug 2025
  5. [5]The state of AI: how organizations are rewiring to capture value · McKinsey · 12 Mar 2025
  6. [6]Usage of AI technologies increasing in EU enterprises · Eurostat · 23 Jan 2025
  7. [7]AI in the Irish workplace: employee survey · Ibec · 2 Sept 2025
  8. [8]Klarna AI assistant handles two-thirds of customer service chats in its first month · Klarna · 27 Feb 2024
  9. [9]Klarna reverses AI push, says customers prefer human support · Forbes · 18 May 2025
  10. [10]Klarna CEO says company will use humans to offer VIP customer service · TechCrunch · 4 Jun 2025
  11. [11]Canaries in the coal mine? Six facts about the recent employment effects of artificial intelligence · Stanford Digital Economy Lab · 26 Aug 2025
  12. [12]Shopify CEO tells employees to prove AI can't do a job before asking for more headcount · CNBC · 7 Apr 2025
  13. [13]The state of AI in 2025: agents, innovation, and transformation · McKinsey · Nov 2025
  14. [14]20% of EU enterprises use AI technologies · Eurostat · 11 Dec 2025
  15. [15]Understanding of AI and its use in Ireland: what CSO data tells us · Central Statistics Office · 20 Aug 2026
  16. [16]2025: the state of generative AI in the enterprise · Menlo Ventures · 9 Dec 2025
  17. [17]No widespread displacement, but the AI employment gap for young workers has widened to 19% · Stanford Digital Economy Lab · 12 Aug 2026

Begin your AI transformation.

Book a call with the founders. Thirty minutes to understand your business, your team, and where AI could actually help. No deck, no pitch.

Talk to us →