Skip to content
ROI & Measurement

Measuring AI ROI: A Practical Framework

A shared methodology finance and engineering can both defend, covering baselining, attribution, and total cost of ownership.

Satori Canton · August 16, 2026 · 14 pages · v2.0

Adoption is a property of the tool. Return is a property of the workflow. Enterprises measure the first because it is easy and hope it implies the second. It does not.

Most large enterprises can now produce a detailed account of AI adoption: seats provisioned, weekly active users, prompts per employee, pilots in flight. Far fewer can produce a defensible account of AI return. The gap is not a reporting problem. It is a measurement design problem, and it has two specific causes: almost nobody captures a baseline of workflow economics before rollout, and almost nobody commits to an attribution method that separates AI-driven change from everything else that changed at the same time.

This paper presents the framework we use to close that gap. It has four parts. First, baseline the workflow before automation, using six measures that take days, not months, to capture. Second, choose an attribution method before rollout: a holdout comparison where the work is divisible, a before-and-after design with an explicit confound ledger where it is not, and a modeled estimate only where effects are too diffuse for either. Third, account for total cost of ownership, including the lines most models omit: inference spend, permanent human review time, rework on AI errors, and ongoing maintenance. Fourth, report the result on one page, in a format finance can audit.

A worked example runs through the paper: an accounts payable automation whose vendor pilot implied roughly $590,000 a year in savings, and whose measured, attributed, fully costed return is a net $332,000 a year. The second number is smaller. It is also real, and it survives scrutiny. That is the trade this framework asks you to make.

Why adoption metrics and return metrics diverge

Adoption metrics (seats provisioned, weekly active users, queries per day) are easy to collect and always trend in a reassuring direction, because usage tends to grow once a tool exists. None of that tells you whether the organization is better off. A support team can roll out an AI assistant to every agent, watch usage climb every week, and still have no defensible answer for whether average handle time improved, whether resolution quality held steady, or whether headcount plans should change as a result.

Return metrics ask a harder, more specific question: for a defined workflow, over a defined period, what changed in cost, time, or quality, and how much of that change is attributable to the AI system rather than to seasonality, a process change that happened alongside it, or a training effect that would have occurred anyway. Answering that question requires instrumentation decisions made before rollout, not after.

Baselining a workflow before automation

A baseline is a single number (or a small set of them) captured before any AI intervention touches the workflow: current cost per unit of work, current cycle time, current error or rework rate. The baseline has to be measured on the same workflow definition that will later be measured post-automation, using the same unit of work, or the comparison is not valid.

The most common baselining failure is not measuring nothing. It is measuring the wrong thing: a team captures overall department cost instead of cost per transaction, or a cycle-time average that mixes work the AI system will never touch with work it will. Both produce a baseline that looks precise and is not comparable to anything measured later.

Attribution: separating AI-driven change from everything else

Once a baseline exists, the harder problem is attribution. If cost per transaction drops 20 percent in the quarter after an AI rollout, and the team also changed its staffing model, adjusted a vendor contract, and shipped an unrelated process improvement in the same window, the 20 percent cannot be assigned to the AI system without more work.

Three attribution approaches, in order of rigor: a holdout comparison (some portion of the workflow continues without AI assistance during the measurement window, so the AI and non-AI groups can be compared directly); a before/after comparison with an explicit list of every other change made in the same period, so confounds are at least visible; and a modeled estimate, used when neither of the above is practical, that states its assumptions plainly enough for someone else to challenge them. The weakest defensible number is still stronger than a number with no stated method behind it.

Total cost of ownership, including inference and rework

The cost side of the ledger is where most ROI claims quietly go wrong, because it is easy to count only the obvious costs (subscription or per-seat fees) and skip the ones that show up later: inference spend that scales with volume, the engineering time spent on prompt and pipeline maintenance, and the cost of rework when AI output requires a full human redo rather than a light edit.

A complete cost model includes: direct inference or API cost per unit of work; the fully-loaded cost of human review time, for workflows that still require it; the fully-loaded cost of rework, weighted by how often a full redo is actually required rather than assumed; and implementation and ongoing maintenance cost, amortized over the measurement period rather than treated as a one-time sunk cost that disappears from the model after launch.

A one-page reporting template finance will accept

The report that survives a budget review is short, states its baseline and attribution method explicitly, and shows the full cost side rather than only the savings side. One page, five sections: the workflow and the unit of work being measured; the baseline, with its measurement date and method; the attribution method used, stated plainly; the full cost model, including inference, review, rework, and maintenance; and the resulting net return, with the specific caveats a skeptical reader would raise. A number without its method attached persuades no one who has been burned by a bad number before.

What’s inside

  • Why adoption metrics and return metrics diverge
  • Baselining a workflow before automation
  • Attribution: separating AI-driven change from everything else
  • Total cost of ownership, including inference and rework
  • A one-page reporting template finance will accept

Who this is for

CFOs and finance leadersHeads of AI or ML platform teamsProgram and portfolio owners

Get the full paper

Sign in to purchase this paper for $49.00.

Sign in

Author

Satori Canton

Founder & Principal

Satori Canton is the founder and principal of ROAI, an advisory practice focused on measuring and improving the return on enterprise AI investment.


Want the numbers behind your own AI investment?

Book a focused session to see where AI creates real economic value in your organization.