Skip to content

Most enterprises measure AI adoption. Almost none measure AI return.

Three widely cited surveys put enterprise AI adoption near ninety percent and measurable financial return under ten percent. The gap is the real story.

Satori Canton

August 14, 2026 · 7 min read

Ask a room of executives whether their company has adopted AI, and nearly all of them will say yes. Ask the same room whether they can point to a specific number, a dollar figure their finance team would sign off on, and the room goes quiet.

That gap between the two questions is not a communication problem. It is the actual state of enterprise AI in 2026, and three independent research efforts have now measured it from different angles and landed on the same conclusion.

The gap between using AI and being changed by it

McKinsey's 2025 global AI survey found that 88 percent of organizations report regular AI use in at least one business function, up from 78 percent the year before. Adoption, by any reasonable definition, is close to universal. But only 39 percent of the same respondents could point to any measurable enterprise-level impact on EBIT, the number a CFO actually cares about. Of that 39 percent, most put the figure below 5 percent. McKinsey classified just 6 percent of respondents as "AI high performers," organizations where AI drives 5 percent or more of EBIT alongside other significant value (McKinsey, "The State of AI," 2025).

MIT's NANDA initiative reached a more severe version of the same finding from a different methodology. Reviewing more than 300 public AI deployments and interviewing 52 organizations, researchers found that despite an estimated 30 to 40 billion dollars in enterprise generative AI spending, 95 percent of organizations saw no measurable return on their pilots. A small number, roughly 5 percent, extracted millions in value from tightly integrated deployments. The rest stayed stuck in pilot purgatory with no path to a number anyone could defend in a budget review (MIT NANDA, "The GenAI Divide: State of AI in Business 2025").

Gartner's forecast, published in mid-2024, predicted the practical consequence of that gap before most of it had played out: at least 30 percent of generative AI projects would be abandoned after proof of concept by the end of 2025, due to unclear business value, escalating costs, and inadequate risk controls (Gartner, July 2024).

SourceWhat it measuredAdoption or useMeasurable financial return
McKinsey, State of AI 2025Global survey of enterprise AI use88% use AI in at least one function39% report any EBIT impact; 6% report 5%+
MIT NANDA, GenAI Divide 2025300+ deployments, 52 org interviewsNear-universal pilot activity5% of pilots show measurable P&L return
Gartner, July 2024 forecastForward-looking predictionN/A30%+ of projects abandoned post-POC by end of 2025

Three different research organizations, three different methods, one consistent shape: adoption is high, and the number of companies that can show their AI spend produced a financial return is small. Not zero. Small.

Why adoption metrics are so easy to report and so easy to misread

Adoption metrics are seductive because they are easy to produce and always trend upward. Seat counts, weekly active users, number of pilots launched: every one of these can be reported honestly and still tell you nothing about whether the company is better off. A team can roll out a coding assistant to a thousand engineers, watch usage climb every week, and still have no answer for whether it changed velocity, defect rates, or headcount planning.

Consider a hypothetical that plays out across most large organizations in some form. A customer-support team adopts an AI drafting tool. Within a quarter, three-quarters of agents are using it daily, and the rollout gets cited in the next board deck as an AI success story. Nobody, at any point, recorded average handle time or first-contact resolution before the tool arrived, so nobody can say whether either number moved. The tool might be saving real money. It might be adding review overhead that cancels out the time it saves. Both stories are consistent with "adoption is high," which is exactly why adoption is the wrong number to report as success.

Return metrics are harder to produce for a specific reason: they require a baseline that existed before the AI tool did, and a decision about which business number the tool is supposed to move. Most organizations skip this step for an organizational reason as much as a technical one. Setting a kill criterion in advance means someone has to put their name on the number that would end the program, and most sponsors would rather keep a pilot alive and ambiguous than own the decision to cancel it. An adoption metric never forces that choice. It only ever goes up, so it never asks anyone to be accountable for anything.

None of these figures mean AI does not work. They mean most organizations have not built the measurement discipline to know whether it works for them specifically. Those are different problems with different fixes, and confusing them leads to the wrong response, more pilots instead of better measurement.

What measuring return actually requires

The organizations in McKinsey's small high-performer group and MIT's small return-generating group were not using fundamentally different models or bigger budgets. What separated them was a measurement practice built before the rollout, not bolted on after. Four things showed up consistently.

A baseline captured before the tool touches the workflow. If nobody records cost, cycle time, or error rate before AI enters a process, there is no honest way to attribute a later change to it. This has to happen in the same week the pilot starts, not the quarter after someone asks for a readout.

A single business metric the tool is accountable to, not a bundle of soft signals. Pick the one number, cost per unit of work, hours saved and redeployed, error rate, that the sponsoring executive would actually cite in a budget conversation. A pilot answerable to five vague goals is answerable to none of them.

A kill or scale decision tied to that number, made in advance. Before the pilot launches, the team agrees on the threshold that triggers scaling and the threshold that triggers stopping. Deciding this after seeing the results is not measurement, it is storytelling.

A review cadence that survives the person who championed the pilot. Programs that depend on one sponsor's memory to get revisited quietly become permanent without ever proving their case. A calendar invite outlives an enthusiastic slide deck.

None of this is complicated. It is also, per MIT's interviews and McKinsey's survey data, what the overwhelming majority of enterprise AI programs are not doing.

What to do Monday

Pick one AI initiative already running in your organization, ideally the one that has consumed the most budget or attention. Ask the team running it a single question: what was the number before, and what is the number now. If nobody can answer that question with a figure rather than a sentiment, that is your actual finding, not a footnote to fix later.

Set the baseline this week for that initiative and for the next one that launches. Name the one business metric it has to move, write down the kill or scale threshold before anyone has seen a result, and put the review on a calendar that does not depend on the person who championed the pilot still being in the room. That is the entire difference between the 88 percent and the 6 percent.

Satori Canton

Founder & Principal

Satori Canton is the founder and principal of ROAI, an advisory practice focused on measuring and improving the return on enterprise AI investment.

Get a second opinion on your AI numbers.