The argument in brief
- The only evidence a finance committee acts on is currency, on history it has already lived. A priced backtest can be produced in about a quarter, and it converts more executives than any presentation.
- Run the policy in parallel before relying on it. The existing process keeps making the real decisions while the policy makes shadow ones, which collects the override signal before anything has been bet.
- Measure the decision, not the forecast. Six measures matter, and conventional planning metrics are largely absent from the list because they describe the instrument rather than the outcome.
- A working system does not fail loudly. It becomes ignorable. Decay is gradual, and it is detectable well before anyone announces it, provided the right two curves are being watched.
- The single most useful leading indicator is the count of surviving spreadsheets, and it is the one nobody wants to publish.
The only evidence that actually funds this
Programmes of this kind are usually justified with a benefits case built on assumptions, benchmarks and a projected improvement in an accuracy metric. Such cases are approved routinely and believed rarely, which is why they tend to be revisited the moment the first budget pressure arrives.
There is a considerably stronger artifact available, and it can be produced before any system is bought. Take the decisions your organisation actually made over the last two years. Price them against the ledger. Then take the same historical data, apply a simple priced policy, and price the decisions it would have made instead. Both numbers are in currency. Both describe events that have already happened. Neither depends on a forecast of the benefit.
Exhibit 1
The same two years, the same data, and the difference expressed in the only unit that funds anything
Three cautions on producing this credibly, each learned the expensive way. Use a policy simple enough that nobody can accuse it of hindsight; a newsvendor calculation against the ledger is usually sufficient and is harder to dispute than anything more sophisticated. Exclude any information that would not have been available at the time of the decision, and be visibly rigorous about it. And present the result as a range, since a single number invites a debate about the third decimal place rather than about the decision.
Running it in parallel before betting on it
A backtest establishes that a better decision was available. It does not establish that the organisation can make one prospectively, and a sensible executive will make that distinction. The bridge between them is a parallel run.
For a defined period, usually one or two planning cycles, the existing process continues to make the real decisions. The policy makes shadow decisions alongside it, recorded but not executed. Each cycle, the two are compared in currency, and every divergence is examined with the planner who would have overruled it.
The parallel run does three things at once, and the second and third are usually undervalued. It produces prospective evidence rather than retrospective evidence, which is what makes the eventual cutover a small decision rather than a large one. It surfaces the override signal described in Paper 04 before anything depends on it, so the first live cycle starts with a model that has already absorbed a season of operator knowledge. And it gives the planning team a low-stakes period in which to disagree with the system and be taken seriously, which does more for adoption than any amount of training.
The one thing to avoid during a parallel run
Scoring the policy on how often it agreed with the planner. Agreement is not the objective, and a policy tuned to maximise it is a policy tuned to reproduce the process you set out to replace. The question is always which decision would have cost less against the ledger.
The six measures that matter
Conventional planning metrics are largely absent from what follows, because they measure the forecast rather than the decision. These six measure the decision.
| Measure | What it tells you | What a bad reading means |
|---|---|---|
| Realised cost of error against the ledger's prediction | Whether the ledger is calibrated | Persistent divergence means the ledger is wrong and everything downstream is mispriced |
| Value added, in currency, against a naive policy | Whether the machinery is earning its keep | Near zero means you have automated a process rather than improved a decision |
| Share of decisions committed without override | Whether the models are trusted | Far below four in five means they are not. At or near 100 percent means overrides are being suppressed, which is worse |
| Cycle time from data close to committed decision | Whether the assembly burden is genuinely gone | A modest reduction means the shadow process is still running |
| Override composition: routine against novel | Whether the system is still learning | See the next section. This is the earliest reliable signal of decay |
| Surviving local master files | Whether the old process has actually been retired | Any number above zero is a live relapse risk, and this is the measure nobody volunteers |
How a working system decays
Systems of this kind rarely fail visibly. Nobody switches them off. They become ignorable, gradually, and the organisation returns to committee without ever having decided to.
The decay is detectable well in advance, and Exhibit 2 shows the two curves that reveal it.
Exhibit 2
Two override curves, and only one of them is supposed to fall
Overrides on routine decisions should decline steadily, because the model is absorbing the reasoning behind them. A routine line that refuses to fall means the models are not trusted, and the cause is almost always that the ranking is not inspectable: planners cannot see why the top option won, so they substitute their own.
Overrides on genuinely novel decisions should stay roughly flat, indefinitely. That line is the system working as designed. A novel line trending toward zero is not maturity; it is judgment being suppressed, usually because someone has begun reporting a blended override rate as a performance metric and the organisation has responded rationally to it.
A single blended override number conceals both failure modes, and is the metric most organisations choose.
The governance that stops the drift
Sustaining the system requires four standing commitments. They are unglamorous, they take perhaps a day a month in total, and their absence is the difference between an architecture that compounds and one that has to be rebuilt in four years.
- Recalibrate the ledger on a fixed schedule. Quarterly is usually right. Compare what the ledger predicted the cost of error would be against what it actually was, and adjust. A ledger that has not been revised in a year is a ledger nobody is using.
- Review model drift explicitly, by name and by owner. Lever-response relationships decay as markets move, and they decay silently. The model steward reports on confidence ranges, not only on central estimates.
- Keep the experiment queue funded. The tests that resolve thin evidence are the only mechanism by which response models improve. They are also the first thing cut in a difficult quarter, which is precisely when the information is most valuable.
- Publish the spreadsheet count. Uncomfortable, trivially cheap, and the single most reliable early warning available. An organisation willing to put that number on a slide each quarter will not drift far without noticing.
This closes the series. The argument across the six papers has been a single one: that recurring operational decisions are a manufacturing problem rather than an analytical one, that they cannot be improved without first being priced, and that the organisation which prices them acquires something more durable than a better forecast. It acquires the ability to see the distribution it has been operating inside the whole time, and to choose its position within it deliberately.