The argument in brief
- A forecast is not a number. It is a distribution over what could happen. The single figure your process circulates is one summary of that shape, and it is not the summary a stocking or staffing decision requires.
- Accuracy metrics select the wrong model. Two models with identical error scores can differ by 30 percent or more in the cost of the decisions they produce, because error metrics are symmetric and the business is not.
- Controllable levers belong inside the model, not as adjustments after it. Price, promotion, marketing and service level move demand. Deciding them first and modelling their effect is the difference between a forecast and a negotiation.
- Where the data is thin, the correct action is to test, not to guess. A system that knows the limits of its own evidence can propose the experiment that would resolve them, which is a capability almost no planning process has.
- The measure of a forecasting function is the cost of the decisions it enables. Reporting accuracy in isolation is reporting on the instrument rather than on the outcome.
What a forecast actually is
Every planning organisation produces a forecast, and almost every one of them produces the wrong object. Not an inaccurate one. The wrong kind of object.
What circulates in the monthly cycle is a single number per item per period. What the business actually faces is a range of possible outcomes with different likelihoods attached. The single number is one summary of that range, usually the mean or something near it, and the act of reducing the range to it discards precisely the information a decision needs.
Exhibit 1
The forecast is a summary of a shape, and it is not the summary a stocking decision needs
Consider what Exhibit 1 shows. The point forecast of eleven units is not wrong; it is a perfectly reasonable summary of that distribution. But a planner deciding how much to stock needs to know something the point forecast cannot express: how bad the bad case is, how likely it is, and what it would cost. Those questions are answerable from the shape and unanswerable from the number.
This is why the most common improvement programme in the discipline, which is to make the point forecast more accurate, so frequently produces disappointing returns. It is an investment in refining a summary that was the wrong summary to begin with.
Why accuracy metrics select the wrong model
The second problem is more serious, because it is actively misleading rather than merely incomplete. The metrics used to choose between forecasting approaches do not measure what the business cares about, and they can and do select the worse option.
Error metrics are symmetric. Being ten units over and ten units under score identically. The business is not symmetric: as Paper 02 established, the two directions carry different costs, land in different budgets, and one is largely invisible. A metric that treats them as equivalent is measuring something the organisation has no economic interest in.
Exhibit 2
Identical accuracy, materially different money: what an error-based bake-off cannot see
Exhibit 2 shows the practical consequence. Two models are indistinguishable on the metric the organisation uses to choose between them, and differ by more than a third in the cost of the decisions they generate. A selection process run on accuracy picks between them by coin flip, and then reports a successful model evaluation.
The organisations that win are not the ones with the smallest forecast error. They are the ones whose errors cost the least.
The remedy is not a better error metric. It is to stop scoring the instrument and start scoring the outcome. Every candidate model should be evaluated by running the decisions it would have produced against the historical record and pricing those decisions with the ledger. That is a harder evaluation to build, it takes a few weeks rather than an afternoon, and it is the only one whose result a finance director has any reason to act on.
Model the drivers rather than adjusting for them
There is a third and quite different failure, and it concerns the things the business controls.
Price changes, promotions, marketing spend, assortment decisions and service commitments all move demand. In most planning processes they are handled as adjustments: a baseline forecast is produced, and then someone uplifts it because a promotion is planned. The uplift is a judgment, it is rarely recorded with its reasoning, and its accuracy is never assessed because the counterfactual is unobservable.
The correct treatment inverts the sequence. The controllable levers are decided first, and the model computes demand conditional on those decisions. The effect of each lever becomes a fitted relationship in its own right, estimated from the organisation's own history of what it did and what followed, and carrying an explicit confidence range rather than a single coefficient.
The practical consequences are worth stating plainly, because they are the reason this matters to a commercial director rather than only to a modeller. Once lever effects are modelled, the question what would happen if we doubled the promotion becomes answerable before the money is committed rather than debated afterwards. The uplift argument that currently consumes part of every consensus meeting becomes a calculation. And the organisation begins accumulating knowledge about its own commercial levers instead of relitigating the same disagreement each year with the same absence of evidence.
When the data is too thin to exploit
Modelling lever effects raises an honest objection: for many decisions, the historical evidence is genuinely weak. The promotion was run once, three years ago, in one region, alongside a competitor action that confounds it. No amount of statistical sophistication conjures information that is not in the record.
This objection is correct, and the response to it is one of the few genuinely novel capabilities the architecture provides. A model that carries explicit confidence ranges knows where its evidence is thin. That knowledge is actionable: it identifies exactly which deliberate experiment would be worth running, and roughly what resolving the uncertainty would be worth in currency.
The organisation therefore gains a queue of proposed tests, ranked by the value of the information they would produce. A controlled price variation in two regions. A holdout on one campaign. A deliberate service-level difference across comparable accounts. Each is small, each is priced, and each converts a permanent argument into a settled fact.
Why this rarely happens without deliberate design
Every commercial function already runs experiments implicitly, in the sense that it does different things in different places for different reasons. What is missing is the discipline of designing a small number of them to be readable afterwards, and an owner whose job includes deciding which uncertainties are worth paying to resolve. Paper 05 gives that role a name and a place in the organisation.
What to ask your analytics team on Monday
This paper is not a brief to rebuild the forecasting function. In most organisations the forecasting function is competent and is being asked the wrong question. Four questions are usually enough to establish where you actually stand.
- Does our forecast carry a distribution, or only a number? If the model produces one internally and the process discards it downstream, the capability already exists and is being thrown away at the interface, which is a much easier problem than it sounds.
- How do we choose between two candidate models? If the answer is an error metric, the selection process cannot see the difference that matters, and the fix is to score candidates on the priced outcome of their decisions instead.
- Are promotions and price changes inputs to the model or adjustments after it? If they are adjustments, the organisation is systematically failing to learn from its own commercial activity.
- Where is our data too thin to trust, and what would it cost to find out? A team that can answer this immediately is further along than most. A team that has never been asked will usually have a good answer within a fortnight.
None of these require a platform decision, and all four are answerable with people already on the payroll. The next paper sets out the machine that these components fit into, and where exactly a human being enters it.