A rollout that promises a benefit spends evidence as well as money. The switch-on plan determines which comparisons survive to the next budget review, so approve it with the business case and ask for four lines: the claim, the comparison and its unit, the measurement and when the comparison window closes.

You are about to approve a rollout of something the business pays for by the seat or by use, and the business case promises faster resolution or fewer errors. The licence covers the whole organisation, so the plan is to switch everyone on in the same fortnight. I have seen organisations on an enterprise-wide licence make that call. The seats are paid for, and every idle one looks like waste.

The plan spends licence money, which the business case prices, and evidence, which it does not. Turn everyone on at once with nothing held back and there is no group left to compare against when the investment comes back for review. Rollout design belongs in the investment decision for that reason, next to the price. A mandatory security control needs no such test, because its case rests on an obligation rather than a promised gain.

In February 2026 the FinOps Foundation changed its mission from advancing the people who manage “the value of cloud” to those who manage “the value of technology”. If you manage technology for its value, ask at approval what you will be able to prove when this comes back.

Cost evidence grows while the comparison runs down

After go-live, the cost evidence gets richer every month. Invoices arrive, seat counts settle and usage logs fill up. The evidence of what would have happened without the tool thins out, and that widening gap is the evidence cost. You retire the old process, and without deliberate staging the people still working without the tool become fewer and less typical of the business. By the third quarter you are defending the value claim with before-and-after figures and assumptions that live in a spreadsheet footnote.

eBay introduced Seller Hub, its analytics dashboard for sellers, at random, which let Bar-Gill, Brynjolfsson and Hak estimate that access to it raised sellers’ revenue by 3.6 per cent on average. Had eBay given every seller the dashboard in the same week, that clean comparison would have gone. Any estimate would then have leaned on stronger assumptions to separate the dashboard from the season.

AI can make the evidence cost higher. At Microsoft, a Copilot field experiment ended in April 2023, about eight months in, because developers in the control group had started seeking access as word of AI coding tools spread. METR found its follow-up developer trial gave an unreliable signal once developers who preferred AI declined to take part and 30 to 50 per cent of participants held back tasks they did not want to do without it. Users may resist going back, and the model and its configuration can change mid-rollout.

I have reviewed seven AI business cases in the past ten months. All seven named the benefit, the cost and the payback date, and none said what the results would be compared against. That leaves each of those organisations more dependent at renewal on whatever observational evidence survives the rollout.

Not every comparison supports the claim you want to make

In Your Cheapest AI Model Might Be Your Most Expensive Decision I costed a successful outcome, and in The AI FinOps Practitioner Is Starting to Look Like an Economist the other ways of producing it. Neither tells you what the tool changed. That takes a comparison, and at a budget review some comparisons bear far more weight than others.

ComparisonWhat it can supportMain caveat
Randomised comparison at the same time, by queue, shift or teamWhat giving that group access changed, over that periodAccess is not use: if some people never use the tool, this measures the effect of offering it
Randomised staggered rollout, with groups going live in waves in a random orderStrong evidence of the effect of access, if the design and analysis are soundMore exposed to changes over time, including the tool changing between waves; the NIH notes a greater risk of bias than a parallel comparison
A non-random rollout compared with the change in a similar untreated group, using a method such as difference-in-differencesA plausible effect, if the groups were comparable and had no reason for their trends to diverge anywayA rollout ordered by who was keen can flatter the early results
Before and after, last year’s figures, or a forecastThat results moved against an expectationDemand, seasonality and every other change that quarter are mixed in
A vendor’s model or users’ estimateAn attributed or perceived benefitWeak evidence of cause. Technical workers surveyed by METR put their median speed gain at 3x and their value gain at 1.4 to 2x; METR’s own 2025 study found people overestimated AI’s effect on their task time by 40 percentage points

Find the smallest unit at which the tool does not leak into the comparison. A single case can be too small in a contact centre, because an agent who learns from tool-assisted calls carries that learning into the untreated ones. A queue or a team may hold. The smaller the valid unit, the closer access can stay to universal while a holdout group survives.

Take the last value report behind a renewal you approved and find its row. Then check whether the original rollout plan could have held back a team or randomised the order, and so moved that report up a row.

Rollout sequencing belongs in the approval

The switch-on plan (who starts first, and whether any group waits) sets how well the benefit claim you sign off can be tested at renewal. So you are approving the plan as well, whether or not anyone shows it to you. The strongest objection is that sequencing is a delivery decision: set the target, let IT move as fast as adoption allows, and measure before and after. I disagree. Measured before and after, the benefit claim is tested on the second-weakest row of the table, by the people who asked for the money.

Speed and evidence are separate questions: a fast rollout randomised by queue or team keeps its comparison, while a slow one ordered by enthusiasm can lose it. In the phased rollouts I have seen, the keenest teams went first. Their early numbers can make the tool look better than it is, because those were the teams most ready to change how they worked anyway.

A Fortune 100 company in the same study as Microsoft’s gave all 3,054 developers in its trial access within two months, but randomised the dates, so some teams started six weeks before others. UW Health randomised 66 providers to receive an AI documentation tool at week 1, 7 or 13, then scaled the programme to more than 600 licences. Staging did not stop either organisation widening access.

Sometimes turning everyone on at once is still the right call, when a competitive window will not wait. That is a fair trade if whoever approves it knows the renewal case will rest on weaker evidence. Staging costs something too: split twelve teams into three equal waves a month apart and a third of them wait one month while another third wait two. If the waves carry equal value and each team’s benefit runs at a steady rate once it has ramped up, each month a team waits is a month of benefit forgone: 12 of the 144 team-months in the first year, or 100,000 dollars on an illustrative benefit of 1.2 million dollars a year. That 100,000 buys a comparison for the renewal, when you commit another year of licence fees (say 500,000 dollars, also illustrative) on whatever evidence survived. Without it, a before-and-after chart may not separate the four-point gain the case promised from a quiet quarter.

What the approval should say

In an illustrative contact centre adding an AI assistant, group twelve teams into balanced blocks by size and call type, and draw the order at random within each block (twelve shows the scheduling; a statistician’s power calculation sets how many teams you need to detect a four-point change). Three waves of four teams go live in early March, April and May. Waves carry more bias risk than a parallel holdout (the table’s second row), but every team has the tool within the quarter, which is easier to approve than holding a group back for a year. If team-level staging is too slow, randomise by queue instead, and write the design into the approval as four lines:

  1. First-contact resolution, the one primary measure, rises by four percentage points (the illustrative target behind the 1.2 million dollars), and customer satisfaction stays above a set floor.
  2. Teams are the unit (or queues, if you randomised those), and until each wave goes live, the teams still waiting are the comparison.
  3. The operations team records both measures weekly for every team from the first week of February, with the model and configuration version in use.
  4. Wave three is the last untreated group, so the comparison window closes when it goes live in early May. The data collected up to then stays usable for the renewal case, even if that is written months later.

Written down before go-live, the claim cannot be quietly redefined by the sponsor or the vendor once the results arrive. If the numbers come in below it, you know by how much and in which teams. I did not hold my own advice to this standard. In January I set termination thresholds in AI Workloads Are Breaking Your FinOps Model, such as 2x returns in 12 months, without saying how anyone would show the return came from the project rather than from everything else that changed that year. A sponsor could meet a threshold like that by choosing the right before-and-after chart, and I should have said so at the time.

The evidence cost of your next rollout is set when you sign its plan, months before anyone opens a dashboard. I would send back any business case that promises a benefit and arrives without the four lines. Before the next one reaches you, ask for them on its first page: the claim, the comparison and its unit, the measurement, and the date the comparison window closes.

Could your next rollout prove its benefit?

Bring one business case waiting for approval, and we will work through its claim, comparison, measurement and closing date with you.

Book a Conversation