IMPORTANT: If you really want to understand this, listen to the audio while you slowly read through the research keeping in synch – trust me.
AI investments deserve ROI cases that survive the boardroom.
A long-form working playbook for C-level leaders — how to construct an AI ROI case that names the total value of change, prices the total cost to serve, and wins the board’s scrutiny, not just its attention.
- Audience
- CEO · CFO · CIO · CTO · COO · CDO/CAIO · Strategy · Board
- Length
- ~16 minutes · 14 sections
- Posture
- Board-ready synthesis
A credible AI ROI case requires more than a vendor deck or a labor-savings spreadsheet. The cases boards approve look like the cases boards can defend. This piece lays out the working playbook — framework, cost stack, measurement architecture, governance, and the 90 days that follow.
Synthesis prepared from McKinsey, Google Cloud, Microsoft, IBM, AWS, Deloitte, NIST, Protiviti, FinOps, and Doctor of DT’s own AI-Enabled Transformation research. Sources cited at the foot of each section and consolidated at the end.
Most AI cases are rejected long before they reach scale — and boards have good reasons.
In IBM’s 2025 C-suite study, only 25% of AI initiatives delivered their expected ROI and only 16% had scaled enterprise-wide. IBM recommends starting with cost-saving use cases, baselining time and quality, and using open orchestration. Google Cloud’s 2,500-executive survey widens the picture: 74% report current ROI, 45% of those reporting productivity gains see at least a doubling, 56% report improved security posture.
Pilots overstate wins. Production undersells cost. Adoption moves slowly. Outcomes are not tied to the metrics the board cares about.
Four reasons boards reject pilot-stage cases
Pilots overstate wins, understate cost
Cherry-picked cohorts, no counterfactual, hand-tuned prompts, and rollout costs billed elsewhere make pilots look better than the production system will.
Value follows a J-curve, not a step
Learning curve, verification tax, and pipeline adaptation leave early ROI negative. Google Cloud calls out realistic financial modelling as the primary lever.
Adoption is the real bottleneck
Workflow redesign, verification, and data readiness slow realisation. Microsoft Worklab reports 79% see AI as crucial, yet 59% struggle to quantify.
The fourth reason is the deepest at the board tier: outcomes are not tied to business metrics. Activity counts — prompts, calls, hours saved — replace value metrics — margin, NPS, risk reduction. Boards discount what cannot be linked to P&L or risk. Protiviti’s board survey makes the engagement side visible: only 26% of boards discuss AI at every meeting; 63% of high-ROI organisations put AI on every agenda; 13% of low-ROI counterparts do.
Sources — IBM 2025 C-suite study · Google Cloud ROI study · Microsoft Worklab · Protiviti global board survey.
AI is an operating-model investment — price it that way and the ROI discussion changes.
Boards fund transformations; they approve tool purchases. The same dollars read differently under different lenses. McKinsey warns against paying a “gen-AI premium” for projects that can be ranked on business case alone.
If a project reads as workflow reinvention — customer operations, marketing & sales, software engineering, R&D, the four functions where McKinsey estimates ~75% of generative-AI value concentrates — the right conversation is portfolio mix and capital allocation. If it reads as a point tool, it is funded inside existing OpEx and the conversation is payback and adoption. Both are valid; confusing them is the most expensive mistake.
McKinsey’s economic potential study places the prize at $2.6T–$4.4T annually. Roughly 75% lives in four functions. Their CFO guide recommends ranking the top 20–30 value-accretive projects regardless of AI tag, picking a small number of high-impact cases, and budgeting for learning.
If a project sits in the four named functions it almost certainly warrants a transformation-grade conversation. If it sits outside, ask whether it is reconstituting an end-to-end workflow, or whether it is one tool in a longer chain. The first case deserves a board conversation; the second belongs in business-unit budgeting.
Source basis — McKinsey’s two references cited above. Heuristic is synthesis.
A credible case captures hard financial value, soft operating value, and strategic option value.
Refuse to collapse the three into “labor savings.” Labor is one line on one tier of three; treating it as the whole case is the most common reason boards discount an ROI submission.
A case that names only efficiency reads as departmental. A case that names all three tiers reads as enterprise.
What shows on the P&L this year.
What compounds before it converts.
What buys the next decade.
Microsoft’s six-dimension value framework tracks closely with this three-tier split. Multi-tier framing is the single most-read board defence.
Eight categories — two routinely forgotten — priced before any revenue line is shown.
Boards recognise under-counted cost faster than they recognise overstated value. A board case that omits two of the eight categories will be sent back; a case that names all eight stands a chance.
| # | Category | What it covers | Share |
|---|---|---|---|
| 01 | Model & API licensing | Foundation-model fees, evaluation licences, sandbox costs. | 15 – 25% |
| 02 | Token & inference | Inference cost per task; non-linear in scale; retries & orchestration overhead. | 10 – 30% |
| 03 | Infrastructure & data | GPU/CPU pools, vector indexes, retrieval pipelines, observability. | 10 – 15% |
| 04 | Integration & UX | Product redesign, workflow rework, last-mile wearable interfaces. | 10 – 15% |
| 05 | Risk & compliance | Bias testing, security review, audit, regulatory preparation. Omitted in most decks. | 5 – 15% |
| 06 | Human & QA | Review loops, exception handling, prompt-ops, model fine-tuning. | 5 – 20% |
| 07 | Change & adoption | Training, communications, internal champions, productivity ramp-down. | 5 – 15% |
| 08 | Lock-in & exit | Switching cost, abstraction layer, replacement runway. Omitted in most decks. | 2 – 10% |
Shares are synthesis. Google Cloud’s chatbot example surfaces the magnitude — roughly $24,910 of estimated monthly value against an estimated $2,700 of monthly cost, payback under two weeks, only when the eight categories are named honestly.
A working playbook executives can defend in a single sitting.
Use as the spine of the business case. Sections F–I unpack steps 1, 6, 7, and the safeguards orgs forget.
Rank the top 20–30 candidates.
Force-rank on value-accretive potential without an AI-tag premium. Pick a small number of high-impact cases.
McKinsey CFO guide · Deloitte balanced portfolio.
Write the baseline down.
Define the do-nothing path and the matched cohort. Freeze baselines before any pilot opens.
IBM, AWS prescriptive guidance.
Map to all three tiers.
Hard, soft, strategic. Each owned by a named executive. Refuse to collapse the case to labor savings.
See section C.
Price the eight-category cost stack.
Risk, change, verification as line items. Boards reward honest accounting and punish omission.
See section D.
Risk-adjusted return, named scenarios.
Apply scenario weights (P25 / P50 / P75). Pair payback and IRR with risk-adjusted ROI; let the board weigh tail risk explicitly.
Synthesis · Deloitte balanced portfolio.
Six metric layers, leading & lagging.
Adoption → productivity → quality → risk → financial → strategic. Cohort tracking, named owner per layer.
See section G · Microsoft, Google, IBM.
Pilot → production → scale, with decision rules and stop conditions.
Each gate has a decision rule, a named owner, a stop condition. Set the next-review cadence before any scale money moves.
AWS decision points · Google J-curve framing.
Pick use cases by value-and-feasibility — with the counterfactual written down.
Two axes. Value-at-stake horizontally, feasibility and readiness vertically. Most disagreements about which projects to scale reduce to which quadrant a project sits in.
Quick proof points
Low risk, fast to deploy. Fund inside existing OpEx. Build conviction and reusable foundations.
End-to-end reimagination
Where McKinsey estimates ~75% of the value lives: customer ops, marketing & sales, software engineering, R&D.
Total prize: $2.6T–$4.4T / yr across four functions.
Defer or descope
Low value, low readiness. Park, or use as cheap learning vehicles.
Capability-build first
High value, low readiness. Treat as multi-quarter investments; build capability, then sequence cases.
Without a counterfactual the case is a memo.
For every shortlisted case, write down three things before the proposal leaves the working group:
What the metric looks like 12 months in if no one does anything. Include the natural improvement the team would have made anyway.
Region, segment, role, or process twin that does not receive AI in the same window.
Per AWS prescriptive guidance: pick autonomy level (assist / recommend / act) and the error budget that goes with it.
McKinsey, IBM, AWS, and Deloitte informed the axes; the choice of axes is synthesis.
Track six metric layers, split leading from lagging, never measure outcomes without the cohort that did not see the system.
Without a counterfactual, every improvement looks attributable to AI — including improvements the team would have shipped anyway.
Adoption and productivity move first; financial lags by one to two quarters.
Boards that ask only about cost and revenue at 90-day review end up killing projects that are tracking. Pair lagging outcomes with leading signals.
Always pair the AI cohort with a matched non-AI cohort.
The cohort is the cheapest line on the budget and the most defensible artefact in the audit trail.
45% of gen-AI users reporting productivity gains saw productivity at least double.
74% of surveyed enterprises report current ROI. 56% report improved security posture.
Microsoft Worklab: time saved alone is insufficient — pair with quality, effort, creativity. Six-layer split is synthesis.
AI is a board-grade risk domain. Boards that treat it that way post higher ROI; boards that do not, post less.
The cadence of the board’s review is itself a financial decision.
The cadence of the board’s review is a financial decision.
Source: Protiviti global board survey.
Name an executive sponsor for every dollar of AI spend.
Without a named sponsor the work routes to committee and never to a number.
NIST AI RMF — govern, map, measure, manage.
Trustworthy-AI controls: bias testing, transparency, accountability. Risk-adjust ROI to make tail events readable.
Establish a quarterly AI portfolio review — not annual.
Score the portfolio like PE firms score deals: kill, double, or hold.
Board materials aligned to audit-committee readability.
Plain-English risk with named consequences. Shadow-AI inventory and stop conditions in the same cover as the wins.
Synthesised from NIST AI RMF, Protiviti global board survey, and Deloitte balanced-portfolio framing.
Five safeguards plus an adoption lever — adapted from the AI-Enabled Transformation research.
Adaptation, in board-readable form, of ideas from Doctor of DT’s AET research.
Value-stream mapping
Reversibility vs consequence
Pre-mortems and red-teaming
Automated circuit breakers
Shadow-AI containment protocol
Cognitive ROI & human value
Adapted from Doctor of DT’s AI-Enabled Transformation framework — doctorofdt.com/ai-enabled-framework-research.
Cost per token is necessary; cost per successful outcome is sufficient.
Token economics is one row of the cost ledger — but its framing quietly determines whether the ROI case ages well or badly.
Cheapest model, shortest prompts, aggressive caching.
Right model for the task — accuracy, retry rate, human review are part of the unit.
A cheaper model that ships wrong answers and forces downstream rework.
A more expensive model that costs less because it solves the problem the first time.
Hidden in plain sight — costs multiply across reasoning chains.
Showback, chargeback, per-task ceilings become part of the case.
Three numbers field research put on the case.
Field reference — FinOps on token economics.
Four sourced reference cases keep the framework close to observable outcomes — not ambitions.
Substitute your own enterprise benchmarks before the board meeting; these are anchors, not targets.
A single chatbot, priced honestly, paid back in under two weeks.
A quarter of initiatives deliver on expected ROI — fewer still scale.
Open orchestration. Baselining first.
Most enterprises now report ROI; the conviction gap is in the data.
Time saved alone is insufficient — six dimensions span value, experience, and risk.
- · Revenue impact
- · Productivity & efficiency
- · Security & risk
- · Employee / customer experience
- · Quality improvement
- · Cost savings
Worklab: 79% see AI as crucial; 59% struggle to quantify.
Sources per case are listed in the References section. Substitute your own benchmark data before the meeting.
Ten items the executive team must have on the table before asking for scale funding.
Treat as decision hygiene, not paperwork. Each row needs a named owner and a signable line.
Cross-references — section E (framework), section G (KPIs), section H (governance), section I (safeguards).
What the next 90 days should look like.
Three phases. Days 1–30 align. Days 31–60 build the case. Days 61–90 present and stage-gate.
Stakeholder map & sponsor
CFO + CIO + line business sign the executive sponsor. Map legal, security, HR, data owners.
Rank the top 20–30 candidates
Score on value, feasibility, autonomy envelope. Force-rank without AI-tag bias.
Inventory shadow AI
Survey + telemetry sweep. Standardise, sandbox, or sunset.
Governance baseline
Map current controls to NIST AI RMF; name owners; agree stop-condition policy.
Top 3 cases: full ROI pack
Apply the seven-step framework. Build the eight-category cost stack and three-tier value.
Cohort & baseline freeze
Match each pilot to a non-AI cohort. Freeze baseline; document the do-nothing path.
Measurement architecture live
Wire the six-layer KPIs. Pair leading and lagging; auto-collect where possible.
Risk-adjusted scenarios
P25 / P50 / P75 explicit. Name kill-switch authority and threshold semantics.
Board pre-read & decision pack
Names, numbers, owners. Decisions on the front page; risks on the back.
Pilot-gate decision
Continue, expand, pause, or kill. Document the decision and the data behind it.
Scale-funding request
Risk-adjusted ROI + payback + NPV pair. Capex shape clearly named.
Cadence & escalation
Quarterly board review booked; audit-committee material staged; escalation path live.
Capacity creation and margin expansion live in the seams between functions.
The 90 days above are where conviction forms — not where transformation completes. Treat the cadence as permanent, not as a project plan.
Quarterly board review · monthly exec stand-up · weekly use-case cadence. If the cohort underperforms, revert. If it leads, double down. Document either decision.
Cadence references — Protiviti board engagement benchmark, NIST governance baseline, AWS stage-gate guidance.
Thirteen primary-source references cited across the article.
Replace with your own enterprise benchmarks before the board meeting — preserve the audit trail, prefer primary sources.
Adapted for section I from Doctor of DT’s AI-Enabled Transformation research.
