The ROI Mandate for AI

IMPORTANT: If you really want to understand this, listen to the audio while you slowly read through the research keeping in synch – trust me.

Skip to content
Doctor of DT · AI research publication

AI investments deserve ROI cases that survive the boardroom.

A long-form working playbook for C-level leaders — how to construct an AI ROI case that names the total value of change, prices the total cost to serve, and wins the board’s scrutiny, not just its attention.

Audience
CEO · CFO · CIO · CTO · COO · CDO/CAIO · Strategy · Board
Length
~16 minutes · 14 sections
Posture
Board-ready synthesis
Doctor of DT · Working playbook

A credible AI ROI case requires more than a vendor deck or a labor-savings spreadsheet. The cases boards approve look like the cases boards can defend. This piece lays out the working playbook — framework, cost stack, measurement architecture, governance, and the 90 days that follow.

Synthesis prepared from McKinsey, Google Cloud, Microsoft, IBM, AWS, Deloitte, NIST, Protiviti, FinOps, and Doctor of DT’s own AI-Enabled Transformation research. Sources cited at the foot of each section and consolidated at the end.

A · The credibility problem

Most AI cases are rejected long before they reach scale — and boards have good reasons.

In IBM’s 2025 C-suite study, only 25% of AI initiatives delivered their expected ROI and only 16% had scaled enterprise-wide. IBM recommends starting with cost-saving use cases, baselining time and quality, and using open orchestration. Google Cloud’s 2,500-executive survey widens the picture: 74% report current ROI, 45% of those reporting productivity gains see at least a doubling, 56% report improved security posture.

25% / 16%
of AI initiatives delivered expected ROI, then scaled enterprise-wide.
45%
of gen-AI users reporting productivity gains saw productivity at least double.
26 / 63 / 13
all boards discuss AI every meeting. High-ROI orgs sit at 63%, low-ROI counterparts at 13%.

Pilots overstate wins. Production undersells cost. Adoption moves slowly. Outcomes are not tied to the metrics the board cares about.

Four reasons boards reject pilot-stage cases

Reason 1

Pilots overstate wins, understate cost

Cherry-picked cohorts, no counterfactual, hand-tuned prompts, and rollout costs billed elsewhere make pilots look better than the production system will.

Reason 2

Value follows a J-curve, not a step

Learning curve, verification tax, and pipeline adaptation leave early ROI negative. Google Cloud calls out realistic financial modelling as the primary lever.

Reason 3

Adoption is the real bottleneck

Workflow redesign, verification, and data readiness slow realisation. Microsoft Worklab reports 79% see AI as crucial, yet 59% struggle to quantify.

The fourth reason is the deepest at the board tier: outcomes are not tied to business metrics. Activity counts — prompts, calls, hours saved — replace value metrics — margin, NPS, risk reduction. Boards discount what cannot be linked to P&L or risk. Protiviti’s board survey makes the engagement side visible: only 26% of boards discuss AI at every meeting; 63% of high-ROI organisations put AI on every agenda; 13% of low-ROI counterparts do.

Sources — IBM 2025 C-suite study · Google Cloud ROI study · Microsoft Worklab · Protiviti global board survey.

B · Reframing the investment question

AI is an operating-model investment — price it that way and the ROI discussion changes.

Boards fund transformations; they approve tool purchases. The same dollars read differently under different lenses. McKinsey warns against paying a “gen-AI premium” for projects that can be ranked on business case alone.

If a project reads as workflow reinvention — customer operations, marketing & sales, software engineering, R&D, the four functions where McKinsey estimates ~75% of generative-AI value concentrates — the right conversation is portfolio mix and capital allocation. If it reads as a point tool, it is funded inside existing OpEx and the conversation is payback and adoption. Both are valid; confusing them is the most expensive mistake.

McKinsey’s economic potential study places the prize at $2.6T–$4.4T annually. Roughly 75% lives in four functions. Their CFO guide recommends ranking the top 20–30 value-accretive projects regardless of AI tag, picking a small number of high-impact cases, and budgeting for learning.

A working heuristic

If a project sits in the four named functions it almost certainly warrants a transformation-grade conversation. If it sits outside, ask whether it is reconstituting an end-to-end workflow, or whether it is one tool in a longer chain. The first case deserves a board conversation; the second belongs in business-unit budgeting.

Source basis — McKinsey’s two references cited above. Heuristic is synthesis.

C · The full value stack

A credible case captures hard financial value, soft operating value, and strategic option value.

Refuse to collapse the three into “labor savings.” Labor is one line on one tier of three; treating it as the whole case is the most common reason boards discount an ROI submission.

A case that names only efficiency reads as departmental. A case that names all three tiers reads as enterprise.

Tier 1 · Hard financial value

What shows on the P&L this year.

Direct cost reduction
Headcount substitution, processing cost, channel cost, vendor consolidation.
Revenue uplift
Conversion lift, win-rate, cross-sell, wallet share.
Margin expansion
Mix shift, premium capture, productivity net of reinvestment.
Working capital
DSO/DPO shifts, inventory turns, capex deferral.
Tier 2 · Soft operating value

What compounds before it converts.

Throughput
Cycle time, queue depth, time-to-decision, jobs per FTE.
Quality & rework
Error rates, first-pass yield, compliance exceptions.
Experience
NPS, CSAT, employee eNPS, ease-of-use.
Speed
Time to insight, response, prototype-to-production cadence.
Tier 3 · Strategic option value

What buys the next decade.

Innovation cadence
New products, time-to-market, experimentation velocity.
Platform leverage
Reusable models, data assets, agent infrastructure.
Resilience
Continuity, supply-chain optionality, scenario coverage.
Data & regulatory
First-party data, audit posture, compliance upside.

Microsoft’s six-dimension value framework tracks closely with this three-tier split. Multi-tier framing is the single most-read board defence.

D · The full cost stack

Eight categories — two routinely forgotten — priced before any revenue line is shown.

Boards recognise under-counted cost faster than they recognise overstated value. A board case that omits two of the eight categories will be sent back; a case that names all eight stands a chance.

#CategoryWhat it coversShare
01Model & API licensingFoundation-model fees, evaluation licences, sandbox costs.15 – 25%
02Token & inferenceInference cost per task; non-linear in scale; retries & orchestration overhead.10 – 30%
03Infrastructure & dataGPU/CPU pools, vector indexes, retrieval pipelines, observability.10 – 15%
04Integration & UXProduct redesign, workflow rework, last-mile wearable interfaces.10 – 15%
05Risk & complianceBias testing, security review, audit, regulatory preparation. Omitted in most decks.5 – 15%
06Human & QAReview loops, exception handling, prompt-ops, model fine-tuning.5 – 20%
07Change & adoptionTraining, communications, internal champions, productivity ramp-down.5 – 15%
08Lock-in & exitSwitching cost, abstraction layer, replacement runway. Omitted in most decks.2 – 10%

Shares are synthesis. Google Cloud’s chatbot example surfaces the magnitude — roughly $24,910 of estimated monthly value against an estimated $2,700 of monthly cost, payback under two weeks, only when the eight categories are named honestly.

E · The framework · seven steps, one spine

A working playbook executives can defend in a single sitting.

Use as the spine of the business case. Sections F–I unpack steps 1, 6, 7, and the safeguards orgs forget.

01 · Use case

Rank the top 20–30 candidates.

Force-rank on value-accretive potential without an AI-tag premium. Pick a small number of high-impact cases.

McKinsey CFO guide · Deloitte balanced portfolio.

02 · Counterfactual

Write the baseline down.

Define the do-nothing path and the matched cohort. Freeze baselines before any pilot opens.

IBM, AWS prescriptive guidance.

03 · Value drivers

Map to all three tiers.

Hard, soft, strategic. Each owned by a named executive. Refuse to collapse the case to labor savings.

See section C.

04 · Full cost

Price the eight-category cost stack.

Risk, change, verification as line items. Boards reward honest accounting and punish omission.

See section D.

05 · Risk-adjusted return

Risk-adjusted return, named scenarios.

Apply scenario weights (P25 / P50 / P75). Pair payback and IRR with risk-adjusted ROI; let the board weigh tail risk explicitly.

Synthesis · Deloitte balanced portfolio.

06 · KPI architecture

Six metric layers, leading & lagging.

Adoption → productivity → quality → risk → financial → strategic. Cohort tracking, named owner per layer.

See section G · Microsoft, Google, IBM.

07 · Stage-gate cadence

Pilot → production → scale, with decision rules and stop conditions.

Each gate has a decision rule, a named owner, a stop condition. Set the next-review cadence before any scale money moves.

AWS decision points · Google J-curve framing.

F · Step 1 of the framework · selecting where to play

Pick use cases by value-and-feasibility — with the counterfactual written down.

Two axes. Value-at-stake horizontally, feasibility and readiness vertically. Most disagreements about which projects to scale reduce to which quadrant a project sits in.

Quadrant A

Quick proof points

Low risk, fast to deploy. Fund inside existing OpEx. Build conviction and reusable foundations.

Quadrant B · the prize

End-to-end reimagination

Where McKinsey estimates ~75% of the value lives: customer ops, marketing & sales, software engineering, R&D.

Total prize: $2.6T–$4.4T / yr across four functions.

Quadrant C

Defer or descope

Low value, low readiness. Park, or use as cheap learning vehicles.

Quadrant D

Capability-build first

High value, low readiness. Treat as multi-quarter investments; build capability, then sequence cases.

Counterfactual discipline

Without a counterfactual the case is a memo.

For every shortlisted case, write down three things before the proposal leaves the working group:

1 · The do-nothing baseline.

What the metric looks like 12 months in if no one does anything. Include the natural improvement the team would have made anyway.

2 · The matched cohort.

Region, segment, role, or process twin that does not receive AI in the same window.

3 · The autonomy envelope.

Per AWS prescriptive guidance: pick autonomy level (assist / recommend / act) and the error budget that goes with it.

McKinsey, IBM, AWS, and Deloitte informed the axes; the choice of axes is synthesis.

G · Step 6 of the framework · measurement architecture

Track six metric layers, split leading from lagging, never measure outcomes without the cohort that did not see the system.

Without a counterfactual, every improvement looks attributable to AI — including improvements the team would have shipped anyway.

Layer 1
Adoption
Active users, weekly depth, abandonment, workflow penetration.
Leading
Layer 2
Productivity
Cycle time, throughput per FTE, tasks closed.
Leading
Layer 3
Quality
First-pass yield, error & rework, exceptions, NPS shifts.
Mixed
Layer 4
Risk
Bias incidents, security findings, audit exceptions.
Lagging
Layer 5
Financial
Cost-to-serve, ARPU / margin, payback, IRR.
Lagging
Layer 6
Strategic
Products shipped, data reuse, optionality, resilience.
Lagging
Leading & lagging

Adoption and productivity move first; financial lags by one to two quarters.

Boards that ask only about cost and revenue at 90-day review end up killing projects that are tracking. Pair lagging outcomes with leading signals.

Counterfactual, again

Always pair the AI cohort with a matched non-AI cohort.

The cohort is the cheapest line on the budget and the most defensible artefact in the audit trail.

Reference benchmark

45% of gen-AI users reporting productivity gains saw productivity at least double.

74% of surveyed enterprises report current ROI. 56% report improved security posture.

Microsoft Worklab: time saved alone is insufficient — pair with quality, effort, creativity. Six-layer split is synthesis.

H · Board oversight · governance as risk surface

AI is a board-grade risk domain. Boards that treat it that way post higher ROI; boards that do not, post less.

The cadence of the board’s review is itself a financial decision.

Engagement correlates with ROI

The cadence of the board’s review is a financial decision.

63%of high-ROI organisations put AI on every board agenda.
13%of low-ROI organisations do.
26%of all boards discuss AI at every meeting — most do not.

Source: Protiviti global board survey.

I · Accountability

Name an executive sponsor for every dollar of AI spend.

Without a named sponsor the work routes to committee and never to a number.

II · Risk controls

NIST AI RMF — govern, map, measure, manage.

Trustworthy-AI controls: bias testing, transparency, accountability. Risk-adjust ROI to make tail events readable.

III · Cadence

Establish a quarterly AI portfolio review — not annual.

Score the portfolio like PE firms score deals: kill, double, or hold.

IV · Disclosure

Board materials aligned to audit-committee readability.

Plain-English risk with named consequences. Shadow-AI inventory and stop conditions in the same cover as the wins.

Synthesised from NIST AI RMF, Protiviti global board survey, and Deloitte balanced-portfolio framing.

I · Field-tested safeguards

Five safeguards plus an adoption lever — adapted from the AI-Enabled Transformation research.

Adaptation, in board-readable form, of ideas from Doctor of DT’s AET research.

Safeguard 1

Value-stream mapping

Map each use case against end-to-end value streams. Capacity creation lives in the seams between functions.
Safeguard 2

Reversibility vs consequence

Reversible, low-consequence moves earn speed; irreversible ones earn stricter governance. Taken from Section 1.5 defined in the Doctor of DT’s AI Enabled Framework Research
Safeguard 3

Pre-mortems and red-teaming

Stress-test the case before each gate. Boards fund evidence the team actually tested the work.
Safeguard 4

Automated circuit breakers

Encode thresholds — accuracy floor, cost ceiling, drift breach, OOD rate — that auto-pause workflows.
Safeguard 5

Shadow-AI containment protocol

Untracked AI use is a parallel ledger most CIOs do not see. Inventory quarterly, route through the same risk gate, decide what to standardise, sandbox, or sunset.
Adoption lever

Cognitive ROI & human value

Pair time-saved with cognitive-load and skill-development metrics — without them, projection beats reality.

Adapted from Doctor of DT’s AI-Enabled Transformation framework — doctorofdt.com/ai-enabled-framework-research.

J · The atomic unit of AI value

Cost per token is necessary; cost per successful outcome is sufficient.

Token economics is one row of the cost ledger — but its framing quietly determines whether the ROI case ages well or badly.

Side-by-side
The narrow lens
What is our cost per token?
The board-ready lens
What is our cost per successful outcome?
Optimisation

Cheapest model, shortest prompts, aggressive caching.

Optimisation

Right model for the task — accuracy, retry rate, human review are part of the unit.

Failure mode

A cheaper model that ships wrong answers and forces downstream rework.

Failure mode

A more expensive model that costs less because it solves the problem the first time.

Multi-agent risk

Hidden in plain sight — costs multiply across reasoning chains.

Multi-agent risk

Showback, chargeback, per-task ceilings become part of the case.

The metrics that matter

Three numbers field research put on the case.

Cost per inference
Combine model, hosting, orchestration. Show per task and per cohort.
Token yield rate
Successful outputs per million tokens consumed. Each retry erodes it.
Inference efficiency index
Value delivered per dollar of inference — the unit the board actually cares about.

Field reference — FinOps on token economics.

K · Sourced anchors

Four sourced reference cases keep the framework close to observable outcomes — not ambitions.

Substitute your own enterprise benchmarks before the board meeting; these are anchors, not targets.

Reference 1 · Google Cloud

A single chatbot, priced honestly, paid back in under two weeks.

≈ $24,910
Estimated value / month
≈ $2,700
Estimated cost / month
< 2 weeks
Payback
Reference 2 · IBM 2025

A quarter of initiatives deliver on expected ROI — fewer still scale.

25%
delivered expected ROI
16%
scaled enterprise-wide

Open orchestration. Baselining first.

Reference 3 · Google · 2,500+ execs

Most enterprises now report ROI; the conviction gap is in the data.

74%
report current ROI
45%
saw productivity ≥ 2×
56%
improved security posture
Reference 4 · Microsoft

Time saved alone is insufficient — six dimensions span value, experience, and risk.

  • · Revenue impact
  • · Productivity & efficiency
  • · Security & risk
  • · Employee / customer experience
  • · Quality improvement
  • · Cost savings

Worklab: 79% see AI as crucial; 59% struggle to quantify.

Sources per case are listed in the References section. Substitute your own benchmark data before the meeting.

L · Pre-funding checklist

Ten items the executive team must have on the table before asking for scale funding.

Treat as decision hygiene, not paperwork. Each row needs a named owner and a signable line.

01
Rank the case among top value-accretive, AI-tag-indifferent.
No “gen-AI premium.” Documented ranking against 20–30 candidates.
02
Map value to all three tiers — not labor alone.
Hard, soft, strategic; each owned by a named executive.
03
Price the full eight-category cost stack.
Governance, change, verification as line items.
04
Define baseline and matched cohort explicitly.
Cohort identified; baseline frozen before pilot.
05
Risk-adjusted return and stop conditions.
P25 / P50 / P75 explicit; named kill authority.
06
Six KPI layers tracked end-to-end.
Adoption → productivity → quality → risk → financial → strategic.
07
Governance — NIST controls + named owners.
Trustworthy-AI controls embedded; shadow-AI inventory quarterly.
08
Architecture flexibility & lock-in analysis.
Optionality priced in; abstraction strategy documented.
09
Cognitive ROI, change, and adoption evidence.
Quality, effort, creativity, eNPS; adoption depth.
10
Board-grade cadence and disclosure standard.
AI on the agenda every quarter; decisions archived.

Cross-references — section E (framework), section G (KPIs), section H (governance), section I (safeguards).

M · Activating the playbook

What the next 90 days should look like.

Three phases. Days 1–30 align. Days 31–60 build the case. Days 61–90 present and stage-gate.

Phase 1Days 1 – 30 · align

Stakeholder map & sponsor

CFO + CIO + line business sign the executive sponsor. Map legal, security, HR, data owners.

Rank the top 20–30 candidates

Score on value, feasibility, autonomy envelope. Force-rank without AI-tag bias.

Inventory shadow AI

Survey + telemetry sweep. Standardise, sandbox, or sunset.

Governance baseline

Map current controls to NIST AI RMF; name owners; agree stop-condition policy.

Phase 2Days 31 – 60 · build the case

Top 3 cases: full ROI pack

Apply the seven-step framework. Build the eight-category cost stack and three-tier value.

Cohort & baseline freeze

Match each pilot to a non-AI cohort. Freeze baseline; document the do-nothing path.

Measurement architecture live

Wire the six-layer KPIs. Pair leading and lagging; auto-collect where possible.

Risk-adjusted scenarios

P25 / P50 / P75 explicit. Name kill-switch authority and threshold semantics.

Phase 3Days 61 – 90 · present & stage-gate

Board pre-read & decision pack

Names, numbers, owners. Decisions on the front page; risks on the back.

Pilot-gate decision

Continue, expand, pause, or kill. Document the decision and the data behind it.

Scale-funding request

Risk-adjusted ROI + payback + NPV pair. Capex shape clearly named.

Cadence & escalation

Quarterly board review booked; audit-committee material staged; escalation path live.

Closing note

Capacity creation and margin expansion live in the seams between functions.

The 90 days above are where conviction forms — not where transformation completes. Treat the cadence as permanent, not as a project plan.

Operating cadence

Quarterly board review · monthly exec stand-up · weekly use-case cadence. If the cohort underperforms, revert. If it leads, double down. Document either decision.

Cadence references — Protiviti board engagement benchmark, NIST governance baseline, AWS stage-gate guidance.

N · References

Thirteen primary-source references cited across the article.

Replace with your own enterprise benchmarks before the board meeting — preserve the audit trail, prefer primary sources.

McKinsey
Economic potential of generative AI
$2.6T–$4.4T annually · ~75% concentrated in four functions.
mckinsey.com / tech-and-ai / the-economic-potential-of-generative-ai
McKinsey · CFOs
Gen AI — a guide for CFOs
Rank top 20–30 projects · avoid paying a gen-AI premium.
mckinsey.com / strategy-and-corporate-finance / gen-ai-a-guide-for-cfos
Google Cloud
How to measure the business value of gen AI
J-curve · learning tax · realistic financial modelling.
cloud.google.com / blog / ai-machine-learning / how-to-measure-the-business-value-of-generative-ai
Google Cloud
Measure the value and impact of your AI
Three-part framework · illustrative chatbot ≈ $24,910/mo value vs ≈ $2,700/mo cost.
cloud.google.com / blog / cost-management / measure-the-value-and-impact-of-your-ai
Google · DORA
Generating value from gen AI — ROI study
2,500+ execs · 74% ROI · 45% 2× productivity · 56% improved security.
cloud.google.com / transform / survey-generating-value-from-generative-ai-roi-study
Microsoft
Measuring the impact of Microsoft 365 Copilot & AI
Microsoft Digital AI Value Framework · six dimensions.
microsoft.com / insidetrack / measuring-the-impact-of-microsoft-365-copilot-and-ai-at-microsoft
Microsoft Worklab
How we measure the value of AI at work
Quality, effort, creativity · 79% see AI as crucial · 59% struggle to quantify.
microsoft.com / worklab / how-we-measure-the-value-of-ai-at-work
IBM
Realize ROI of AI agents
25% deliver expected ROI · 16% scale enterprise-wide · open-orchestration advice.
ibm.com / think / insights / realize-roi-ai-agents
AWS
Agentic AI economics — measuring success
Baseline, autonomy level, break-even analysis, decision points.
docs.aws.amazon.com / prescriptive-guidance / latest / agentic-ai-economics
Deloitte
AI tech investment ROI
Balanced digital portfolio · shared-value scorecard.
deloitte.com / insights / digital-transformation / ai-tech-investment-roi
Protiviti
AI in board-meeting discussions — global survey
26% boards every meeting · 63% high-ROI orgs on every agenda.
protiviti.com / us-en / ai-board-meeting-discussions-global-survey
NIST
AI Risk Management Framework
Trustworthy AI built into design, development, use, evaluation.
nist.gov / itl / ai-risk-management-framework
FinOps Foundation
Token economics — the atomic unit of AI value
Cost per inference · token yield · showback / chargeback.
finops.org / insights / token-economics-the-atomic-unit-of-ai-value

Adapted for section I from Doctor of DT’s AI-Enabled Transformation research.

Doctor of DT · Working playbook

Build the case, defend it, scale it.

The cases boards approve look like the cases boards can defend. Capture the total value of change, price the total cost to serve, hold a real counterfactual, and update the case quarter after quarter.