Jon Moshier / Notes / Reference Class Forecasting budding
Note · From the Notebook

Reference Class Forecasting

How forecasting from a class of similar past projects corrects the planning fallacy, where the method breaks down on choosing the class, and what it cannot fix.

Reference class forecasting predicts a project’s outcome not from its own plan but from the recorded outcomes of a class of similar projects already finished. It is the operational form of the “outside view.” Kahneman and Tversky named the underlying idea in 1979; Flyvbjerg turned it into a method governments now mandate. The broader Decision Frameworks note treats it as one debiasing tool among several. This note is about the method itself: how it works, why picking the class is the hard part, and the line between what it corrects and what it cannot touch.

The method is three steps, and the first is the trap

The procedure is short. Identify a reference class of past, comparable projects. Build the distribution of their actual outcomes, cost overruns, delays, demand shortfalls. Place your project in that distribution and read off the prediction, and adjust it sparingly. Tetlock’s superforecasters do the same thing under a different name, “comparison classes,” and the discipline that separates them is starting with the base rate before touching case-specific detail. The habit is a single question: how often do things of this sort happen in situations of this sort?

Kahneman’s own demonstration is the cleanest. Writing a high-school curriculum in Israel, his team estimated two years to a finished draft. He then asked the team’s curriculum expert to recall how long comparable teams had taken from the same stage: no fewer than seven years, with about 40% never finishing, and Kahneman’s team rated below average. They pressed on anyway. The project took eight years, and the curriculum was never used. The inside view, built from the specifics of this team and this plan, was confidently wrong. The outside view, built from the dull record of similar teams, was right and ignored.

The reference class problem is the whole difficulty

Defining “a project of this sort” is not mechanical. Choose a class too narrow and there is no statistical signal; too broad and the cases are not comparable. Philosophers call this the reference class problem: finite data yields probabilities only over groupings, and no rule uniquely fixes the grouping. In forecasting practice it acquires a sharper name, “reference class tennis,” the back-and-forth where two parties each insist on the class that flatters their position. A sponsor wants their bridge compared only to recent bridges with the same procurement model; a skeptic wants it in the pool of all bridges. There is no view from nowhere that settles it.

This is the load-bearing weakness. The method’s accuracy is downstream of a judgment call that the method does not make for you. Identifying the right class weighs similarities and differences across many variables and decides which ones drive the outcome, which is judgment, not arithmetic. Good practice is to err wide. A larger, slightly less similar class beats a tiny, perfectly similar one, because the planning fallacy lives in the optimism of the inside view, and almost any honest outside view dampens it.

Institutionalized as a number you must add

The strongest evidence for the bias the method targets is its skew. Across hundreds of transport projects nine in ten run over budget, rail overrunning by about 45% on average, with demand overestimated by roughly 51%. Honest noise would scatter around zero; this leans one way, every time, which is the fingerprint of systematic error rather than bad luck. Flyvbjerg’s later survey of more than 16,000 projects found that only 0.5% land on budget, on time, and on benefits.

Governments responded by converting the outside view into a mandatory line item. In 2003 the UK Treasury’s Green Book required appraisers to add an explicit “optimism bias uplift” to cost estimates, and in 2004 the Department for Transport, with Flyvbjerg and COWI, set those uplifts empirically by reference class. The supplementary guidance assigns a category and an upper-bound adjustment to each: standard civil engineering carries an uplift of up to 44%, non-standard civil engineering up to 66%, with the figure dialed down only as specific risks are demonstrably mitigated. The forecaster no longer argues about whether to be pessimistic; the default pessimism is supplied, and the burden shifts to justifying any reduction.

What it corrects, and the part it cannot reach

A forecast is not a decision. The Edinburgh tram was among the first projects to use reference class forecasting, in 2004, and still came in around £300m over budget and three years late. A method that produces an honest number changes nothing if the organization overrides the number.

There is a deeper limit. Reference class forecasting is built to cure the planning fallacy, an honest cognitive error. But Flyvbjerg argues much overrun is not error at all. It is, on this account, strategic misrepresentation: planners and promoters lowball cost and inflate benefit on purpose, because the optimistic forecast is what gets the project approved and funded. No debiasing tool touches a deliberate lie, and worse, a forecaster who knows the “right” outside-view number can still report the wrong one if the incentive points that way. The same overrun demands opposite cures depending on cause, a fooled brain versus a rigged bid, and the recent literature warns that the popular story overstates how universal optimism bias is: the evidence is mixed and some projects finish under estimate. Reference class forecasting is the right tool for the first problem and close to useless against the second.

Try it

Build your own reference class (1-2 hours, any spreadsheet). Pull your last 8 to 10 finished tasks where you recorded a time or cost estimate. Put the estimate beside the actual and divide. Take the median ratio: that is your personal planning-fallacy multiplier, and for most people it sits above 1.3. Now look at the spread, not just the center. Honest noise would scatter the ratios around 1.0; yours will cluster above it. That one-way skew, in your own data, is the same fingerprint Flyvbjerg measured across a nation, and it is what tells you the error is bias rather than chance.

Play reference class tennis on a real estimate (30 minutes). Take a forecast you care about and write down three different reference classes for it, ordered from narrowest to broadest. Find or estimate the base rate for each. Watch the prediction move as the class widens, and notice which class you instinctively reached for first. The gap between the flattering narrow class and the sobering broad one is the judgment the method cannot make for you.

See also

Sources

← All notes Read recent essays →