Stop estimating your project from its own details. Reference class forecasting reads the answer off the projects that already finished - and the evidence says it wins.
Reference class forecasting is the practice of estimating a project not from its own details but from the recorded outcomes of comparable projects that have already finished. It is the most rigorously validated cure we have for the tendency of plans to come in late, over budget and under-delivered - and it works by a deceptively simple move: instead of asking your team how long this will take, you ask how long projects like this one actually took. The forecast stops being an opinion and becomes a reading off history.
The method has an unusually strong pedigree for a planning technique. Its theoretical foundation was laid by Daniel Kahneman and Amos Tversky in the 1970s, in the work on judgement under uncertainty that won Kahneman the Nobel Memorial Prize in economics in 2002. It was turned into an operational procedure by Bent Flyvbjerg, who set out its three steps in the Project Management Journal in 2006. It was endorsed by the American Planning Association in 2005, adopted by HM Treasury and the UK Department for Transport in 2004, and has since become the standard against which optimistic business cases are checked in British public procurement. It is, in other words, not a management fashion. It is the closest thing project planning has to an evidence-based standard of care.
This guide covers the whole of it: what reference class forecasting is and what makes it different from every other estimating technique; the evidence that inside-view forecasting fails, and fails in a specific, predictable direction; the intellectual history from the planning fallacy to the Nobel to the Green Book; how to execute all three steps in practice, including how to build a reference class when no public database exists for your industry; how to convert a distribution into an uplift you can actually put in a budget; the failure modes and the honest objections; and how the method combines with Monte Carlo simulation and modern decision intelligence tooling rather than competing with it. Every figure quoted is linked to the source it came from; every number used to illustrate a hypothetical is labelled illustrative.
Reference class forecasting is a debiasing method. That is the first and most important thing to understand about it, because it explains every design choice that follows. It is not a better model, a smarter algorithm, or a more granular work breakdown. It does not attempt to improve the quality of your team’s judgement. It attempts to remove your team’s judgement from the load-bearing part of the forecast altogether, and to replace it with something judgement cannot distort: the record of what actually happened to projects like yours.
The reason for that design is a finding that has held up across five decades of research. Errors in forecasting are not random. If they were random, they would cancel out - some projects would come in as far under budget as others came in over, and the average error across a large sample would sit near zero. That is not what the data shows. Flyvbjerg tested exactly this proposition against the available evidence and found that the distributions of forecasting error are, in his words, “consistently and significantly non-normal with averages that are significantly different from zero.” The conclusion he drew is the premise of the whole method: the problem is bias, not inaccuracy as such. And bias cannot be fixed by trying harder, because trying harder is done by the same biased mind.
The distinction at the heart of reference class forecasting is Kahneman and Tversky’s contrast between two ways of thinking about a future event. The inside view approaches a project by immersing itself in the specifics: the scope, the team, the technology, the dependencies, the known obstacles. It builds a mental model of how the work will unfold, extrapolates that model forward, and produces an estimate. It is the natural, intuitive and near-universal way to plan. When a project manager decomposes a programme into tasks, assigns durations, sums them, adds a contingency and presents a date, that is the inside view in its purest operational form.
The outside view refuses to look at the project at all, at least at first. It treats the project as an instance of a category - a light rail scheme, an ERP migration, a warehouse fit-out, a clinical trial, a product launch in an adjacent market - and it asks what happened to the other members of that category. It does not try to predict which specific things will go wrong. It observes that in projects of this type, things go wrong at a certain rate and with a certain magnitude, and it assumes, absent strong evidence to the contrary, that this project will be no different. Flyvbjerg describes the statistical content of the move precisely: reference class forecasting “consists of regressing forecasters’ best guess toward the average of the reference class and expanding their estimate of credible interval toward the corresponding interval for the class.”
What makes the outside view powerful is not that it knows more than the inside view. In one obvious sense it knows dramatically less - it deliberately discards almost everything specific about the project. What it has instead is immunity. Because the forecaster is not required to construct scenarios, imagine failure modes, or assess their own team’s competence, they cannot get those things wrong. As Flyvbjerg puts it, in the outside view “project managers and forecasters are not required to make scenarios, imagine events, or gauge their own and others’ levels of ability and control, so they cannot get all these things wrong. Human bias is bypassed.” The reference class implicitly contains every cause of overrun that has ever affected projects of this type, including the ones nobody on your team has heard of, the ones that had not been invented yet, and the ones that were unknowable at the time each of those projects was approved.
This is also why reference class forecasting resists a criticism that sinks most debiasing techniques. Kahneman and Tversky found that awareness of a cognitive illusion does not dissolve it - knowing that you are prone to optimism does not make you less optimistic, any more than knowing about an optical illusion makes the lines look equal. Debiasing methods that rely on the estimator noticing and correcting their own bias therefore tend not to work. Reference class forecasting does not ask anyone to notice anything. It changes what data the forecast is made from, which is a structural fix rather than a psychological one.
The operational core of the method is short enough to state in full. Flyvbjerg specifies that reference class forecasting for a particular project requires three steps: identification of a relevant reference class of past, similar projects, where “the class must be broad enough to be statistically meaningful but narrow enough to be truly comparable with the specific project”; establishing a probability distribution for the selected reference class, which “requires access to credible, empirical data for a sufficient number of projects within the reference class to make statistically meaningful conclusions”; and comparing the specific project with the reference class distribution in order to establish the most likely outcome for the specific project.
Each of these steps is harder in practice than it looks on the page, and later sections of this guide take them one at a time. But notice what is absent from the list. There is no step in which you enumerate the risks to your project. There is no step in which you assess the strength of your team, the quality of your requirements, or the sophistication of your governance. There is no step in which you argue about whether the estimate feels right. The method is deliberately austere, because every place where judgement enters is a place where bias can re-enter with it.
Notice, too, what the output is. A reference class forecast does not produce a number; it produces a distribution, and then a number chosen from that distribution at a level of confidence you have consciously selected. This is the same shift from point estimate to distribution that underpins all probabilistic forecasting, and it has the same consequence: contingency stops being a round-number convention and becomes a priced decision about how much certainty you are buying.
It is not benchmarking. Benchmarking compares your project’s planned parameters against typical planned parameters - cost per square metre, cost per kilometre of track, function points per developer-month - and asks whether your plan is in line with peers. That is a useful sanity check on the plan, but it compares estimate to estimate. Reference class forecasting compares your estimate to other projects’ outcomes, which is a categorically different comparison and the only one that captures systematic bias. If every comparable project overran by forty per cent, a plan that benchmarks perfectly against their plans is a plan that will overrun by forty per cent.
It is not a risk register. A risk register is an inventory of the things a team has thought of, scored by likelihood and impact and summed into a reserve. It is inside-view by construction: it can only contain the risks somebody imagined, quantified with numbers somebody supplied. The expert report prepared for the Edinburgh Tram Inquiry by Flyvbjerg and Alexander Budzier makes the point bluntly, observing that the result of conventional risk assessment “looks objective and quantitative, even scientific. In reality it is subjective and qualitative, based on judgement.” Reference class forecasting is the outside check on that judgement, not a substitute for the discipline of identifying risks - you still want to know what might go wrong so you can manage it. It is a substitute for believing the register’s total.
It is not a Monte Carlo simulation, though the two are frequently confused and work beautifully together. A Monte Carlo simulation is a computational engine: it samples repeatedly from probability distributions you have supplied for each uncertain input and reports the resulting distribution of outcomes. Its rigour is real, but it is entirely downstream of the ranges you feed it. If those ranges come from the same optimistic team that produced the point estimate, the simulation will faithfully compute the distribution of an optimistic model. Reference class forecasting is about where the ranges come from. Used together - outside-view calibration of the inputs, Monte Carlo to combine them - you get the best of both, and later in this guide we set out exactly how the combination works.
Finally, it is not a way of predicting what will go wrong on your project. This disappoints people, because a forecast that tells you a number without telling you a story feels unsatisfying. But the refusal to tell a story is the point. Every attempt to explain, in advance, exactly how a project will overrun is an exercise of the same imagination that produced the optimistic estimate in the first place. The reference class knows that projects like yours overrun by a given amount without knowing why yours will, and the historical record suggests it is right far more often than the story is.
The case for reference class forecasting rests on an empirical claim that is easy to state and, by now, extraordinarily well documented: forecasts made from the inside view are wrong in a consistent direction, by large margins, across every domain where anyone has bothered to keep score, and they have not improved. If that claim were false - if inside-view estimating were merely noisy - the sensible response would be to average more estimates and move on. It is not noise. It is bias, and this section lays out the record.
The most systematically studied domain is transport infrastructure, because governments build a lot of it and, unusually, publish both what they planned to spend and what they actually spent. The picture from that data is not ambiguous. In the analysis underpinning reference class forecasting, Flyvbjerg reports that for transportation infrastructure projects, inaccuracy in cost forecasts measured in constant prices averages 44.7% for rail, 33.8% for bridges and tunnels, and 20.4% for roads. These are averages, not worst cases: the typical rail project costs about forty-five per cent more than the figure used to approve it.
44.7% - average cost forecast inaccuracy for rail projects in constant prices, against 33.8% for bridges and tunnels and 20.4% for roads. Over the seventy-year period for which cost data are available, accuracy in cost forecasts has not improved.
The finding about the seventy-year period deserves its own emphasis, because it is the single most damaging fact for the technical explanation of forecasting error. Over seven decades, computing power grew by orders of magnitude, estimating methods professionalised, project management became a discipline with certifications and bodies of knowledge, and the volume of historical data available to estimators expanded enormously. Accuracy did not improve. As Flyvbjerg notes, if imperfect data and models were the main explanation for inaccuracy, “one would expect an improvement in accuracy over time, since in a professional setting errors and their sources would be recognized and addressed.” Substantial resources went into better data and better models. It changed nothing, which tells you the problem was never data or models.
The independent study by Flyvbjerg, Skamris Holm and Buhl, covering 258 transport infrastructure projects, found the same pattern from a different angle: cost escalation averaged around twenty-eight per cent across the sample, with rail worst, fixed links in the middle and roads least affected. Their paper carries a title that has become famous in the field - Underestimating Costs in Public Works Projects: Error or Lie? - and the question in it is not rhetorical, as we will see.
Cost is only half the business case. The other half is demand: how many passengers, users, customers or transactions the thing will attract once built. Here the record is, if anything, more alarming, and it fails in the direction that flatters the project. Flyvbjerg reports that average inaccuracy for rail passenger forecasts is −51.4%, with 84% of all rail projects wrong by more than plus or minus twenty per cent. Read that carefully: the typical rail project attracted roughly half the passengers its business case promised. For roads, average inaccuracy in traffic forecasts is 9.5%, with half of all road forecasts wrong by more than plus or minus twenty per cent - less biased on average, but still wildly imprecise at the level of the individual scheme.
−51.4% - average inaccuracy in rail passenger demand forecasts, with 84% of rail projects wrong by more than ±20%. Over the thirty-year period for which demand data are available, accuracy in rail and road traffic forecasts has not improved.
The compounding effect is what makes this lethal to decision-making. A business case is typically a ratio: benefits over costs. If costs are underestimated by forty-five per cent and benefits are overestimated by roughly double, the resulting benefit-cost ratio is not slightly optimistic, it is wrong by a factor. Flyvbjerg describes the consequence as “inaccuracy to the second degree,” noting that benefit-cost ratios are “often wrong, not only by a few percent but by several factors.” Every downstream appraisal that depends on those two numbers - economic viability, environmental assessment, affordability - inherits the distortion. The organisation is not making a slightly optimistic decision. It is making a decision on numbers that bear no reliable relationship to what will happen.
It would be convenient if this were a quirk of civil engineering. It is not. Flyvbjerg is explicit that comparative research shows the problems, causes and cures identified for transportation apply to “a wide range of other project types including concert halls, museums, sports arenas, exhibit and convention centers, urban renewal, power plants, dams, water projects, IT systems, oil and gas extraction projects, aerospace projects, new production plants, and the development of new products and new markets.”
Information technology is the domain most business readers will recognise from personal experience, and it has been studied closely. McKinsey, working with the University of Oxford, examined 5,400 large IT projects and found that on average they ran 45% over budget and 7% over time while delivering 56% less value than predicted - and that 17% of large IT projects go so badly they threaten the existence of the company. In the Harvard Business Review, Flyvbjerg and Budzier reported on a sample of around 1,500 IT projects with an average cost overrun of 27%, but with a tail that dominates the average: one project in six was a “black swan” with a cost overrun averaging around 200% and a schedule overrun of nearly 70%.
The construction sector tells the same story from the practitioner’s side. KPMG’s Global Construction Survey found that only 31% of respondents’ projects came within 10% of budget in the preceding three years - a figure reported by the people running the projects, not by an external critic. McKinsey’s work on construction’s digital future found large projects across asset classes typically taking 20% longer to finish than scheduled and running up to 80% over budget. And Flyvbjerg’s summary of the whole body of evidence, which he calls the Iron Law of Megaprojects, is that cost overrun is the norm rather than the exception across megaprojects, with roughly nine in ten exceeding their budgets.
The breadth matters because it defeats the most common local objection. Every industry believes its own overruns are explained by conditions peculiar to it - the planning system, the labour market, the regulator, the vendors. When the same signed, sized bias appears in railways, hospitals, software, films, dams and Olympic Games, the explanation cannot be industry-specific. Something common to how humans and organisations produce forecasts is doing the work. Our guide to why projects fail goes deeper into the mechanics; the point here is only that the failure is general.
If the error is not technical, what is it? Flyvbjerg tested three families of explanation - technical, psychological and political - and found that the last two carry the weight. The psychological explanation is optimism bias: a cognitive predisposition, found in most people, to judge future events in a more positive light than actual experience warrants. Under optimism bias, forecasters make honest mistakes. They believe their numbers. Our dedicated guide to optimism bias in business covers the mechanism in detail.
The political explanation is strategic misrepresentation: forecasters and managers deliberately overestimate benefits and underestimate costs in order to increase the likelihood that their project, and not the competition’s, gains approval and funding. This is not a fringe hypothesis; it is the reading that Flyvbjerg, Skamris Holm and Buhl argued their data supports, and it explains features of the record that optimism alone cannot - in particular why the bias fails to decay with experience among professionals who are repeatedly and publicly proven wrong.
Flyvbjerg draws the distinction sharply: “Optimism bias and strategic misrepresentation are both deception, but where the latter is intentional, i.e., lying, the first is not, optimism bias is self-deception.” The two are not rivals but complements, and which dominates depends on context. Where political and organisational pressures are low, optimism bias explains most of the error; where competition for scarce funds is fierce, strategic misrepresentation takes over. In a large organisation, both are usually present at once - an honestly optimistic estimate, produced by a team that also knows what number will get approved.
This matters enormously for what you should expect reference class forecasting to achieve, and it is the subject of a later section on barriers. In brief: where the cause is honest optimism, the method is welcomed, because nobody objects to a technique that makes their forecasts better. Where the cause is strategic, the method is resisted, because its whole purpose is to remove the discretion that made the strategic number possible. As Flyvbjerg notes, in that second situation “the demand for accuracy is simply not there - and barriers are high.” A better forecasting method cannot, by itself, fix an incentive problem.
One more feature of the evidence shapes how reference classes must be built. The distribution of overruns is not merely shifted to the right of zero; in several important domains it is not shaped like a normal distribution at all. In a study of 5,392 IT projects published in the Journal of Management Information Systems, Flyvbjerg, Budzier, Lee, Keil, Lunn and Bester found that IT project cost overruns follow a power-law distribution - many projects with relatively small overruns, and a fat tail containing a smaller number with extreme ones.
5,392 IT projects - analysed to test the shape of the cost-overrun distribution. Overruns follow a power law, not a normal distribution: if managers assume normality, “they may be unwittingly exposing their organizations to extreme risk by severely underestimating the probability of large cost overruns.”
Source: Flyvbjerg, Budzier, Lee, Keil, Lunn & Bester (2022), Journal of Management Information Systems
The practical consequence is severe and counterintuitive. Under a normal distribution, the mean and the mode sit close together, extremes are vanishingly rare, and a reserve set at a modest multiple of the standard deviation buys near-certainty. Under a power law, none of that holds: the mean is dragged upward by the tail, the extremes are rare but not rare enough to ignore, and a reserve sized by normal-distribution intuition is systematically too small. It also means the projects that ruin organisations are not drawn from a different population than the ordinary ones. They are the same population, sampled from further out.
This has a direct methodological implication that we return to when building reference classes: do not remove the disasters. The instinct to strip outliers from a dataset is a good one in many statistical contexts and a catastrophic one here, because the outliers are the risk. The Edinburgh Tram expert report makes the case explicitly, warning that “a common misconception is that Black Swans are freak occurrences to be excluded from reference classes,” and noting that extreme projects are typically not caused by exotic catastrophes such as disease outbreaks or terrorism but by “multiple adverse events occurring simultaneously” - which is to say, by ordinary things co-occurring, exactly the situation your reference class exists to price.
Reference class forecasting has an unusually traceable intellectual history, and it is worth knowing, both because it establishes the method’s credentials and because the original examples explain the mechanism better than any abstract description. The line runs from a pair of psychologists studying prediction errors in the 1970s, through a Nobel prize in 2002, to a British Treasury directive in 2003 and an operational forecasting method in 2004 - a rare case of a laboratory finding becoming mandatory government practice within a working lifetime.
The theoretical and methodological foundations of reference class forecasting were first described by Daniel Kahneman and Amos Tversky in 1979, in a paper on intuitive prediction and corrective procedures published alongside the more famous prospect theory work of the same year. Their broader research programme established three findings that between them make the case for a structural rather than a psychological remedy: that errors of judgement are often systematic and predictable rather than random, manifesting bias rather than confusion; that many errors of judgement are shared by experts and laypeople alike; and that errors remain compelling even when one is fully aware of their nature.
That third finding is the one that forecloses the obvious alternative. If awareness cured bias, the remedy for optimistic estimates would be training. It does not. As Kahneman and Tversky put it, awareness of a perceptual or cognitive illusion does not by itself produce a more accurate perception of reality. What awareness can do, they argued, is enable one to identify situations in which the normal faith in one’s impressions must be suspended and in which judgement should be controlled by a more critical evaluation of the evidence. Reference class forecasting is a method for exactly that critical evaluation - you notice you are in a situation where your impressions cannot be trusted, and then you stop using them.
Kahneman and Tversky located the specific mechanism in the treatment of distributional information. Human judgement, they found, is generally optimistic because of overconfidence and “insufficient regard to distributional information” - meaning the record of how similar things have turned out. People underestimate the costs, completion times and risks of planned actions and overestimate the benefits of the same actions. Dan Lovallo and Kahneman later named this pattern the planning fallacy and traced it to actors taking an inside view, focusing on the constituents of the specific planned action rather than on the outcomes of similar actions already completed. Our guide to the planning fallacy covers the phenomenon in its own right.
The prescription follows directly from the diagnosis, and Flyvbjerg singles it out as the pivotal advice in the entire literature. Kahneman and Tversky wrote that “the analysts should therefore make every effort to frame the forecasting problem so as to facilitate utilizing all the distributional information that is available.” Flyvbjerg comments that this “may be considered the single most important piece of advice regarding how to increase accuracy in forecasting through improved methods.” Everything reference class forecasting does is an implementation of that one sentence.
Kahneman’s own account of how he discovered the effect remains the clearest illustration available, and it has the advantage of being about him. Some years before the theory was formalised, Kahneman was part of a team of academics and teachers developing a curriculum for a new subject area for high schools in Israel. In time the team began to discuss how long the project would take. Everyone wrote an estimate on a slip of paper. The estimates ranged from 18 to 30 months - a reasonable spread, produced by people with genuine expertise, all of whom were looking at the same work.
Then one team member posed a different question to a distinguished expert in curriculum development who was in the room: recall as many projects similar to ours as you can, think of them as they were at a stage comparable to ours, and ask how long it took them to reach completion. After a while, the expert answered with some discomfort. Not all the comparable teams he could think of had ever completed their task - about 40 per cent of them eventually gave up. Of those that did finish, he could not think of any that completed in less than seven years, nor any that took more than ten.
Asked whether the present team was more skilled than the earlier ones, the expert said no; his impression was that it was slightly below average in resources and potential. The two forecasts came from the same person, about the same project, minutes apart: eighteen to thirty months, and seven to ten years with a forty per cent chance of never finishing at all. The wise decision at that point, Kahneman later reflected, would probably have been to disband. Instead the team ignored the pessimistic information and carried on. They completed the project eight years later, and their efforts went largely wasted - the resulting curriculum was rarely used.
Every element of the method is visible in that story. The inside view produced a coherent, expert, badly wrong estimate. The outside view - the same expert, asked to enumerate a reference class and report its distribution - produced an estimate that turned out to be almost exactly right. And the group, presented with both, chose the one it preferred, which is the behavioural obstacle that no amount of methodological rigour eliminates.
The contrast between inside and outside views has been confirmed experimentally. In one study Flyvbjerg cites, a group of students enrolling at a college were asked to rate their future academic performance relative to their peers in their major. On average these students expected to perform better than 84% of their peers, which is of course logically impossible; the forecasts were biased by overconfidence. A second group of incoming students from the same major were first asked about their entrance scores and their peers’ scores, and only then about their expected performance. That single diversion into outside-view information - information both groups already had - reduced the second group’s average expected performance ratings by 20%. Still overconfident, but substantially more realistic.
Note how small the intervention was. Nobody was trained, warned, incentivised or corrected. They were simply asked a question that made distributional information salient before they made their estimate. That is the entire ergonomics of reference class forecasting: it is less a calculation than a change in the order in which questions are asked.
Kahneman was awarded the Nobel Memorial Prize in Economic Sciences in 2002 for the body of work on decision-making under uncertainty that this method rests on. Tversky, who had died in 1996, was acknowledged in the announcement but could not share the prize, as the Royal Swedish Academy of Sciences does not award it posthumously. Three years later the method crossed from psychology into professional planning practice. In April 2005, on the basis of a study of inaccuracy in demand forecasts for public works projects, the American Planning Association officially endorsed reference class forecasting, recommending that “planners should never rely solely on conventional forecasting techniques.”
April 2005 - the American Planning Association endorsed the method: “APA encourages planners to use reference class forecasting in addition to traditional methods as a way to improve accuracy. The reference class forecasting method is beneficial for non-routine projects … Planners should never rely solely on civil engineering technology as a way to generate project forecasts.”
Source: American Planning Association (2005), quoted in Flyvbjerg (2006)
The wording of that endorsement repays attention on two points. First, the APA framed reference class forecasting as an addition to traditional methods, not a replacement - the outside view is a check on the inside view, and organisations that abandon detailed planning in favour of a percentile have misunderstood the instruction. Second, it singled out non-routine projects as where the method is most beneficial, which is precisely the class of project where teams are most convinced that history does not apply to them.
The decisive institutional step came from the British Treasury. The 2003 revision of the Green Book, the UK government’s appraisal and evaluation guidance, identified for large public procurement what its supplementary guidance calls “a demonstrated, systematic, tendency for project appraisers to be overly optimistic. To redress this tendency appraisers should make explicit, empirically based adjustments to the estimates of a project’s costs, benefits, and duration.” Crucially, it specified where those adjustments should come from: “it is recommended that these adjustments be based on data from past projects or similar projects elsewhere.” That sentence is reference class forecasting written into government policy.
The Treasury attached teeth. Departments were told that in future, the allocation of funds for large public procurement would depend on valid adjustments for optimism, and were encouraged to collect the data needed to inform those adjustments where none existed. In response, the UK Department for Transport commissioned Flyvbjerg, in association with COWI, to build the methodology for transport - producing empirically based optimism bias uplifts for selected reference classes and guidance on applying them. The work was carried out in 2003 and 2004 and published by the Department in August 2004; from that date, local authorities applying for transport funding were required to take optimism bias into account using those uplifts. Reference class forecasting had gone from a psychology paper to a condition of funding in twenty-five years.
The first step is the one that decides whether the whole exercise is worth anything, and it is where most attempts go wrong. The requirement, in Flyvbjerg’s formulation, is that the class “must be broad enough to be statistically meaningful but narrow enough to be truly comparable with the specific project.” Those two demands pull in opposite directions, and managing the tension is the craft of the method. Too narrow and you have four projects and no distribution; too broad and you have a distribution that describes something other than your project.
The most common error is to select on features that feel comparable rather than features that predict overrun. A team building a warehouse management system will instinctively look for other warehouse management systems, on the grounds that the domain is the same. But domain is rarely the dominant driver of cost risk. Scale, novelty to the organisation, the number of interfaces to existing systems, the procurement route, the number of stakeholders with veto power and the maturity of the requirements at the point of approval typically matter far more than whether the software manages warehouses or invoices.
The Edinburgh Tram expert report puts the principle in explicitly statistical terms and gives a worked illustration of the trap. Discussing whether a project should build its own bespoke class, it advises that “if projects construct their own reference class, statistical analysis should be used to decide which project types to include” - and then observes that “the light rail projects above are statistically similar to other rail projects and therefore a reference class only of light rail projects would make the error of discarding valuable information.” In other words, a class restricted to the most superficially similar projects was demonstrably worse than a broader one, because the narrowing bought no additional comparability while costing a great deal of statistical power.
The practical test is therefore not “does this project look like mine?” but “is there evidence that projects with this attribute overrun differently from projects without it?” If two subclasses have statistically indistinguishable overrun distributions, splitting them is pure loss. If they differ, keep them apart. Where you have enough data to test this, test it. Where you do not - which in a private company is most of the time - err toward the broader class and compensate by choosing a higher percentile, rather than pretending to a precision your sample cannot support.
The transport uplifts offer a useful model of the process at scale. The types of scheme under the Department’s responsibility were divided into distinct categories where statistical tests, benchmarking and other analyses showed that the risk of cost overrun within each category could be treated as statistically similar. A reference class of completed comparable projects was then established for each category. The guidance was built from a sample of 260 transport projects, including a reference class of 46 rail projects; the road distribution reported by Flyvbjerg was built from 172 completed and comparable projects; and the light rail cost risk analysis in the Edinburgh work drew on 63 historic light rail projects.
260 transport projects - formed the sample behind the UK Department for Transport guidance, including a reference class of 46 rail projects. The uplifts recommended for rail projects, including light rail, were P50 = 40% and P80 = 57%.
Source: Flyvbjerg & Budzier (2018), Report for the Edinburgh Tram Inquiry
Two features of that design are worth stealing. The first is that the categories were derived from the data rather than from organisational convenience: schemes were grouped because their overrun risk behaved similarly, not because they sat under the same directorate. The second is that the class sizes vary considerably - 172 for roads, 46 for rail - and the method tolerates that, because what matters is whether the distribution has stabilised, not whether it has hit a round number.
Most organisations are not the Department for Transport. If you run a mid-sized business deciding whether to open a second location, take on debt or launch a product, there is no published reference class for what you are about to do, and the absence is the usual reason the method never gets tried. That conclusion is too quick. There are four sources of distributional information available to almost any organisation, and using any of them beats using none.
A minimum viable reference class is smaller than most people assume and more useful than none. The honest framing is that with a small class you are not measuring a percentile precisely; you are establishing an order of magnitude for the correction and, above all, establishing that the correction is not zero. Moving a team from “we will hit the date” to “projects like ours have historically landed somewhere between thirty and eighty per cent over, so we should plan for the middle of that and fund the top of it” is the bulk of the available value, and it does not require a hundred data points.
The single most damaging thing you can do when assembling a class is to make it tidy. Every organisation contains a strong social pressure to exclude the embarrassing projects - the one where the sponsor left, the one where the vendor collapsed, the one everyone agrees was a special case. Those exclusions systematically remove the right tail, which is the part of the distribution the whole exercise exists to measure. As we saw above, the extreme projects are generally not caused by exotic events but by several ordinary adverse events landing at once. Special-casing them is equivalent to assuming that ordinary adverse events will not co-occur on your project, which is precisely the assumption the record contradicts.
Apply one rule and defend it: a project enters the class if it met the inclusion criteria at the time it was approved, regardless of how it turned out. Decide the criteria first, then pull the projects. If you find yourself constructing an argument for why a particular disaster does not belong, you are almost certainly rediscovering the reason your organisation’s estimates have historically been low.
With a class defined, the second step is to establish the probability distribution of its outcomes. This sounds like the mechanical part, and conceptually it is - you are computing a set of percentiles from a sample. But the measurement choices you make here determine whether the resulting numbers mean anything, and three of them cause almost all the trouble: what you measure the outcome against, whether you adjust for inflation, and what shape you assume.
A cost overrun is a ratio, and the denominator is a choice. The same project can be described as forty per cent over budget or ten per cent over, depending on whether you measure against the earliest feasibility figure, the outline business case, the final business case, or the contract award. This is not a technicality; it is the most common way a reference class forecast is quietly rendered meaningless, because a distribution built on one baseline and applied to an estimate at another baseline is measuring two different things.
The convention in the published work is to measure against the budget at the time of the decision to build - in the UK, the point at which the business case is presented and the go or no-go is given. Flyvbjerg is explicit that the established uplifts “should be applied to estimated budgets at the time of decision to build a project.” The Edinburgh Tram expert report flags, as a defect in the official guidance, that this was not always respected in practice: the underlying studies “measure cost overruns based on the final decision to build (i.e. the final business case),” while the transport appraisal guidance “uses those numbers as uplifts for the outline business case stage.” Since estimates at outline stage are less mature and therefore more optimistic, applying a final-business-case uplift there understates the correction.
The practical rule is simple to state and requires discipline to follow: the baseline of your distribution and the baseline of the estimate you are uplifting must be the same stage of maturity. If you only have outcome data measured from a later, firmer baseline, and you are applying it to an early, softer estimate, you must uplift by more than the class suggests - and you should say so explicitly rather than let the mismatch pass silently into the budget.
The headline transport figures are reported in constant prices for a reason: an overrun measured in nominal terms conflates project failure with general inflation, and in a high-inflation period this can make a well-run project look disastrous or, worse, provide a ready excuse for one that was not. Strip inflation out of both the baseline and the outcome before computing the ratio. Where the class spans many years or several currencies, this matters more than any refinement to the percentile calculation.
The same like-for-like discipline applies to scope. If a project delivered eighty per cent of what was promised at ninety per cent of the budget, recording it as a ten per cent underrun is a lie of omission - and it is exactly how many public projects end up described as broadly on budget. Edinburgh is the canonical case: the scheme that was eventually delivered was substantially shorter than the one that was approved. Where scope was cut, either normalise the outcome to the original scope or record the scope reduction alongside the cost figure so that the class carries the truth. A reference class that silently rewards descoping will teach your organisation to descope.
Finally, decide up front whether you are building a class for cost, for schedule, or for benefits, and keep them separate. They have different distributions and different biases - the transport data shows costs overrunning by tens of per cent while demand forecasts were wrong by half - and a single blended “project performance” figure obscures both. In practice the three are also used for different decisions: cost drives funding, schedule drives commitments to customers, and benefits drive whether the project should be done at all.
Once assembled, the distribution of overruns in a reference class almost never looks like a bell curve. It is bounded on the left - a project cannot come in more than a hundred per cent under budget - and unbounded on the right, which alone guarantees a right skew. In domains such as IT the skew is severe enough to be a power law, as the study of 5,392 projects established. Even in the transport data, where the shape is less extreme, the Edinburgh report notes that the real light rail curve “differs from the idealized S-curve” that a symmetric distribution would produce.
This makes the arithmetic mean an actively misleading summary. In a right-skewed distribution the mean sits above the median, dragged up by the tail, so it overstates the typical project; yet it simultaneously understates the risk, because it says nothing about how far out the tail extends. Neither the reassuring reading nor the alarming one is correct. The right summary is the set of percentiles - and, in particular, the percentile that matches the level of certainty you actually need, which is the subject of the next section.
Report the distribution as a cumulative curve wherever you can. The Edinburgh report describes the standard presentation: the level of certainty on one axis and the cost risk on the other, where “P50 means that the forecast is 50% certain and has thus a 50% likelihood of being exceeded, P80 means that the forecast is 80% certain and has a 20% likelihood of being exceeded.” A cumulative curve answers the only question a decision-maker really has - how much do I need to hold to be this sure? - without requiring them to interpret a histogram. The same logic drives the S-curves in our guide to probabilistic forecasting and the outputs of Incertive’s probability distribution view.
One last discipline: record the sample size and the date range alongside the percentiles, and revisit them. A reference class is a living object. Projects complete, the class grows, and the distribution shifts - and an organisation that keeps scoring its own outcomes gets a class that improves every year while its competitors are still arguing about whether the last one was a special case.
The third step converts the distribution into a number you can put in a budget. This is where reference class forecasting differs most visibly from every other estimating technique, because the output is not a single answer but a menu: a schedule of uplifts, each attached to a level of certainty, from which the organisation must consciously choose. The choice is not a technical one. It is a statement of risk appetite, and it belongs to whoever is accountable for the money.
The mechanics are straightforward. Having established the distribution of overruns in the class, you read off the percentile corresponding to the confidence you want, and that percentile is the uplift you apply to your estimate. The lower the risk of overrun you are willing to accept, the higher the uplift. The published UK transport figures make the trade-off concrete. For a road project, a willingness to accept a 50% risk of cost overrun requires an uplift of 15%; accepting only a 10% risk requires 45%. For rail, the same two levels require 40% and 68% respectively. At the P80 level - a one-in-five chance of exceeding the budget - the uplifts are 32% for roads and 57% for rail.
Three things fall out of that table that no inside-view estimate can tell you. First, the cost of certainty is quantified: moving a road scheme from a coin-flip to a one-in-ten chance of overrun costs thirty percentage points of budget, and that is now a decision someone can take deliberately rather than a number someone chose because it felt prudent. Second, the shape differs by class - Flyvbjerg notes that risk reduction becomes increasingly expensive for roads and fixed links below 20% risk, while for rail the cost of increased risk reduction rises more slowly, albeit from a much higher level. Third, and most usefully in an argument, the difference between road and rail at the same confidence level is not a matter of opinion. It is what the two classes did.
Which percentile you should use depends on a structural question about the project, not on how confident anyone feels. Flyvbjerg’s guidance is precise about the distinction. The 50% percentile is pertinent to the investor with a large project portfolio, where cost overruns on one project may be offset by cost savings on another; funding every project at P50 across a big enough portfolio should, in aggregate, roughly balance out. The upper percentiles of 80 to 90 per cent should be used when investors want a high degree of certainty that cost overrun will not occur, for instance in stand-alone projects with no access to additional funds beyond the approved budget.
That framing is the single most useful thing to take from this section into a real budgeting conversation, because it converts an argument about optimism into an argument about structure. If your organisation is running forty projects a year and can genuinely move money between them, P50 is defensible and P80 across the portfolio is wasteful. If this is the one big bet, there is no offsetting portfolio and the P50 figure is a coin flip on the company. The UK Department for Transport, for its part, typically accepted a 20% risk of overrun - the P80 level - for large investments in local transport infrastructure.
The same question is worth asking about which *kind* of failure you are protecting against. Cost, schedule and benefits have separate distributions and different consequences: a stand-alone project might reasonably be funded at P80 on cost while committing externally to a P50 date, or the reverse, depending on whether money or the calendar is the binding constraint. Our guide to project risk tolerance works through how to set these levels deliberately rather than by default.
Flyvbjerg gives two examples that show exactly how the arithmetic lands. In the first, a group of project managers preparing the business case for a new motorway decide that the risk of cost overrun must be less than 20%. They therefore apply an uplift of 32% to their estimated capital expenditure budget: an initially estimated £100 million becomes £132 million. Had they instead decided that a 50% risk of overrun was acceptable, the uplift would have been 15% and the final budget £115 million.
In the second, project managers preparing the business case for a metro rail project decide they want 80% certainty of staying within budget. The rail uplift at that level is 57%, so an initial capital expenditure budget of £300 million becomes £504 million. At 50% certainty the final budget would have been £420 million. The gap between those two numbers - eighty-four million pounds - is the price of moving from a coin flip to a one-in-five chance of overrun on a single scheme, and the point of the method is that somebody now has to look at that number and decide.
Notice what the second example implies about the honesty of the original figure. The team’s own estimate was £300 million. The reference class says that a rail project which is 80% likely to stay within budget must be funded at £504 million. Nothing about the project changed between those two numbers; only the source of the forecast did. That gap is the size of the bias, and in a domain like rail it is not a rounding error but two-thirds of the original budget again.
The obvious response from any competent team is that they are better than the class - better managed, better scoped, better contracted. Sometimes that is true, and the method allows for it, but the burden of proof is deliberately steep. Flyvbjerg’s formulation is that “only if project managers have evidence to substantiate that they would be significantly better at estimating costs for the project at hand than their colleagues were for the projects in the reference class would the managers be justified in using lower uplifts.” The symmetric case also applies: if there is evidence that the managers are worse at estimating than their colleagues, higher uplifts should be used.
The word doing the work there is evidence. Confidence is not evidence. A better methodology is not evidence unless there is data showing that projects using it overran less. The relevant evidence is track record: if your organisation has run twelve projects of this type and consistently landed nearer the good end of the class distribution, that is a real reason to sit lower in the class, and it is also something you can only know if you have been scoring your own outcomes - which is why the calibration habit described later in this guide is not a nicety but the precondition for ever earning a discount.
The Edinburgh expert report is pointed about how this discretion gets abused in practice, recommending that “in practice, downward adjustments to risk and optimism bias uplifts ought to pass a critical test of objectivity to be justified.” In the Edinburgh case specifically, Ove Arup concluded that “the justification for reduced Department for Transport optimism bias uplifts would appear to be weak,” because the advanced risk analysis that might have warranted a reduction had not been carried out. The lesson generalises: an uplift reduced by assertion is an uplift deleted.
The published cases are all public infrastructure, so it is worth walking the same arithmetic through a commercial decision. Suppose a mid-sized distributor is deciding whether to open a second warehouse, and the internal estimate for fit-out and systems is £2.4 million with a nine-month timeline (illustrative). Step one: the class is not “second warehouses” - there are none - but the organisation’s last eleven capital projects above £500,000, plus four comparable fit-outs described in industry data. Step two: measured against the figure approved at sign-off, in constant prices, that class shows a median overrun of 22% and a P80 of 61% (illustrative).
Step three is the conversation that matters. This is a stand-alone commitment; there is no portfolio of other warehouses whose underspend could absorb an overrun, and the company cannot raise another tranche of capital quickly. That argues for P80, which means funding £3.86 million rather than £2.4 million (illustrative). The board may well decide that at that price the project is not worth doing - and if so, reference class forecasting has just done its job, because the alternative was discovering the same thing eighteen months later with the money spent. That is the uncomfortable half of the method that its advocates should say out loud: an honest outside view kills projects, and the projects it kills are disproportionately the ones that would have failed.
Britain is the natural case study for reference class forecasting because it is the jurisdiction that adopted the method first, applied it at national scale, and then - in the Edinburgh Tram - produced the test case in which the outside view and the inside view disagreed sharply, were both recorded, and can now be scored against what actually happened. Very few management techniques get a natural experiment this clean.
The Treasury’s supplementary Green Book guidance on optimism bias sets out generic adjustment ranges to be used, in its words, “in the absence of more robust evidence,” derived from a study by Mott MacDonald into the size and causes of cost and time overruns in past projects. The table is short enough to be worth reproducing in full, because it is the most widely used ready-made reference class in existence and because the magnitudes surprise people who have not seen it. The upper-bound capital expenditure uplifts are:
200% - the upper-bound optimism bias uplift on capital expenditure for equipment and development projects - the category that covers software and systems development - against 44% for standard civil engineering and 24% for standard buildings.
Source: HM Treasury, Supplementary Green Book Guidance: Optimism Bias
The 200% figure for equipment and development is the one that stops people, and it should. That category covers software and systems development - the work most modern organisations do most of. The British government’s official, evidence-derived starting assumption for an ICT development project is that it may cost three times the estimate. Set against the McKinsey-Oxford finding of 45% average overrun on large IT projects and the power-law tail identified across 5,392 projects, the number stops looking eccentric and starts looking like an honest read of a badly-behaved distribution.
The procedural design around that table is as instructive as the numbers. The guidance sets out five steps, and the second is the one that changes behaviour: always start with the upper bound. Appraisers are told to “use the appropriate upper bound value for optimism bias from Table 1 above as the starting value,” and only then to reduce it “according to the extent to which the contributory factors have been managed,” expressed as a mitigation factor between 0.0 and 1.0. Ideally, the guidance says, optimism bias should be reduced to its lower bound before contract award - but reductions must be justified: “clear and tangible evidence of the mitigation of contributory factors must be observed, and should be independently verified, before reductions in optimism bias are made.”
This is a deliberate inversion of the normal burden of proof, and it is the single most transferable idea in the whole document. In conventional practice, the estimate starts optimistic and contingency must be argued for, which places the burden on whoever is worried. Here the estimate starts pessimistic and the discount must be argued for, which places the burden on whoever is confident. The difference in outcome is enormous, because in most organisations the person who has to make the argument loses.
The guidance also notes something that is easy to miss and important in practice: the upper-bound percentages “relate to the average historic optimism bias found at the outline business case stage for traditionally procured projects,” and higher adjustments may therefore be required at an earlier stage in the appraisal process. Earlier means less certain, and less certain means a bigger correction, not a smaller one - the opposite of the way early-stage estimates are usually treated, where a rough number is quietly given the benefit of the doubt.
The first recorded practical use of the new uplifts came in October 2004, in the planning of the Edinburgh Tram. Ove Arup and Partners Scotland had been appointed by the Scottish Parliament’s Edinburgh Tram Bill Committee to review the business case for Line 2, developed on behalf of the promoter, Transport Initiatives Edinburgh. The promoter’s business case estimated a base cost of £255 million with an additional allowance for contingency and optimism bias of £64 million - about 25% - giving total capital costs of approximately £320 million.
Arup applied the Department for Transport uplifts to the base cost and calculated the 80th percentile value for total capital costs - the value at which the likelihood of staying within budget is 80% - at £400 million, being £255 million multiplied by 1.57. The 50th percentile came out at £357 million. Arup further remarked that these figures were likely to be conservative, because the Department recommends applying its uplifts at the time of the decision to build, and Line 2 had not yet even reached outline business case stage, meaning risks and corresponding uplifts would be substantially higher. Their conclusion was that “it is considered that current optimism bias uplifts may have been underestimated.”
So the record, in 2004, was unambiguous. The promoter said roughly £320 million. The outside view said £400 million at P80, and probably more. The gap was not a matter of interpretation; it was the difference between an allowance chosen by the people who wanted the project built and a percentile read off forty-six comparable rail schemes.
What happened next is documented at length in the expert report Flyvbjerg and Budzier prepared for the Edinburgh Tram Inquiry. The final business case for the tram forecast the cost of Phase 1a, airport to Newhaven, at £498 million, stating that the estimate included a risk adjustment expected to be a P90 estimate - against a funding commitment of up to £500 million from the Scottish Government and £45 million from the City of Edinburgh Council. In other words, the project told its funders it was ninety per cent certain of landing inside the Scottish Government’s cap, with roughly two million pounds to spare on that line alone. The report records that in real terms, the Edinburgh Tram’s eventual cost overrun was +52%. The line that opened was substantially shorter than the one that had been approved.
+52% - the real-terms cost overrun of the Edinburgh Tram - against a final business case whose estimate was presented as a P90 figure, and against a 2004 reference class forecast that had put the P80 cost roughly 25% above the promoter’s own total.
Source: Flyvbjerg & Budzier (2018), Report for the Edinburgh Tram Inquiry
The Scottish Government’s response to Lord Hardie’s final inquiry report, published in September 2023, records that ten causes of failure were identified and that government funding remained capped at the agreed £500 million throughout. The overspend, in other words, landed on the city.
The most consequential conclusion in the expert report is not about Edinburgh at all. Assessing why a project with a conventional, competent risk management regime still got its risks so badly wrong, the report finds that the tram “established a risk management regime (systems, tools, processes) that was in line with typical risk management regimes of UK infrastructure projects at the time,” and that this regime, “despite the best intentions, is not getting risks right.” It goes further, concluding that inside-view quantitative risk analysis - a risk register plus a Monte Carlo simulation - is insufficient to generate extreme downside scenarios and that, “by creating a false sense of certainty, may add risk instead of reducing it, as appears to have been the case in Edinburgh.”
That is a strong claim and worth restating plainly: a rigorous-looking quantitative risk process, run on inside-view inputs, can leave an organisation worse off than no analysis at all, because it converts an uneasy guess into a confident number. The Edinburgh final business case did not lack quantification. It had a risk register, a Monte Carlo simulation and a figure labelled P90. What it lacked was any outside check on whether the inputs to all that machinery bore any relation to how comparable schemes had actually turned out.
The report’s third finding closes the loop: “optimism in expert reviewers is difficult to root out, unless all analyses are based on hard, empirical data.” Independent review, the usual institutional answer to an over-optimistic business case, is not sufficient on its own, because reviewers are subject to the same biases and are frequently anchored by the numbers they are reviewing. What makes review effective is giving the reviewer a reference class to review against.
Reference class forecasting does not arrive in an empty field. Every organisation already has a way of dealing with estimating uncertainty, and the method has to be understood in relation to those incumbents - partly to know when it adds value, and partly because most of them remain useful and the honest position is that the outside view sits alongside them rather than sweeping them away.
The overwhelming majority of business plans handle uncertainty with a round-number buffer: add ten per cent, add a month, add a quarter. This is reference class forecasting’s most common competitor and its weakest. A conventional contingency has three defects. It is not derived from anything, so it cannot be defended when challenged and is usually the first thing cut. It is uniform across project types, so a routine office fit-out and a first-of-a-kind systems integration get the same buffer despite class distributions that differ by an order of magnitude. And it is invariably too small: a ten per cent contingency against a class whose median overrun is thirty per cent is not prudence, it is a rounding error with a governance process attached.
The comparison with the Treasury table is instructive here. A ten per cent contingency would be below the *lower* bound for non-standard civil engineering and roughly a fifth of the upper bound for standard civil engineering. Against the equipment and development category it is not in the same universe. Whatever else can be said about conventional contingency, it is not calibrated to anything, and the moment you place it next to a real class distribution the arbitrariness is obvious to everyone in the room - which is, in practice, one of the fastest ways to get an organisation to take the method seriously.
The risk register is the professional standard and it is genuinely valuable, but for a different job than the one it is usually asked to do. Its strength is management: it names specific threats, assigns owners, tracks mitigations and creates accountability for the things a team can actually influence. Its weakness is measurement. As the Edinburgh report describes the conventional process, impacts and likelihoods are “quantified, typically using subjective best guesses and rarely based on hard empirical data from past projects,” and the register total is then treated as the risk estimate.
The structural limitation is that a register is a list of known unknowns. It cannot contain the risk nobody thought of, and the historical record is emphatic that unlisted causes account for a large share of overrun - which is precisely why a reference class, which contains the effects of every cause including the unlisted ones, produces bigger numbers than a register does. When your register total and your reference class forecast disagree, the class is not being pessimistic. It is including things your workshop did not.
The right relationship is complementary and directional: use the reference class to set the size of the reserve, and use the register to manage the project down within it. Keeping both, and treating the gap between them as a measure of how much you have not thought of, is more informative than either alone. Our guide to quantifying business risk works through the practical mechanics of moving from a register to a quantified model.
This is the comparison most often got wrong, because the two techniques operate at different layers and are not substitutes at all. Monte Carlo simulation is a method for combining uncertainties: given distributions for each input, it samples them thousands of times and reports the distribution of the result. It is the correct tool for that job and there is no serious alternative to it. What it does not do - what it cannot do - is tell you whether the input distributions are honest. Feed it optimistic ranges and it will return, with great precision, the distribution of an optimistic model.
Reference class forecasting operates one layer up. It is a method for establishing what the inputs, or the aggregate output, should look like based on evidence rather than judgement. Used together, the sequence is: derive the ranges or the overall correction from the class; feed those into the simulation; use the simulation to combine them, model correlations and produce the S-curve. The Edinburgh report describes the same pairing from the other direction, noting that a Monte Carlo simulation can account for correlations between risks and is the industry standard for modelling the full range of futures, while insisting that the inside-view inputs must be checked against the outside view for the result to mean anything.
The failure mode to watch for is a simulation whose output is narrower than the reference class distribution. If your model says there is a 90% chance of landing within 15% of budget, and the class of comparable completed projects shows a median overrun of 30%, the model is not more precise than history - it is wrong, and the false sense of certainty it creates is exactly the pathology the Edinburgh report identified. Treat the reference class as an external validity check on any simulation you run.
Three-point estimation asks estimators for an optimistic, most likely and pessimistic value and combines them, usually with a PERT weighting. It is a genuine improvement on a single number, because it forces the estimator to acknowledge a range, and it is the natural on-ramp to probabilistic thinking for a team that has never done any. But it remains an inside-view technique in outside-view clothing: all three points come from the same judgement, and a pessimistic case supplied by an optimistic estimator is, by construction, not pessimistic enough. The college-entrant experiment described earlier makes the mechanism visible: asked to rate themselves without outside-view information, students placed themselves above 84% of their peers, which cannot be true of a whole cohort. A “worst case” produced by the same faculty inherits the same distortion, and it inherits it in the tail - the one region where being wrong is expensive.
The productive combination is to anchor the three points on class data rather than intuition: let the reference class supply the pessimistic tail, and let the team’s knowledge of the specific project supply the detail in the middle. That preserves what the inside view is genuinely good at - knowing what the work consists of - while removing its authority over the part it is reliably bad at, which is the shape of the downside.
Benchmarking compares plans to plans and therefore inherits any bias common to the industry; if everyone estimates optimistically, benchmarking certifies your optimism as normal. It answers a real and useful question - are we planning to spend an unusual amount per unit? - but it is not an outside view of outcomes and should never be presented as one.
Independent expert review is the other standard institutional check, and Flyvbjerg is careful about its limits. Outside experts are not immune to cognitive bias: they can be anchored by the team’s figures, they may lack the project-specific knowledge to assess it properly, and they carry their own overconfidence. Their value rises sharply when they are given a structured framework - a reference class and a distribution - to review against, rather than being asked for a professional opinion on a spreadsheet. The Edinburgh finding that optimism in expert reviewers is difficult to root out unless analyses rest on hard empirical data is the sharpest available statement of this point.
And versus doing nothing at all - which remains, quietly, the most widely practised alternative - the case does not need making. The comparison that matters is between an organisation that knows the distribution of its own past outcomes and one that does not, and the second is making every capital decision blind while believing it is being careful. Our comparison of structured analysis against gut feeling sets out what changes when that flips.
Almost everything written about reference class forecasting is written about megaprojects, and that is a problem for the overwhelming majority of organisations that will never build a tram line. The megaproject literature exists because governments publish their numbers, not because the method is only valid at that scale - the underlying psychology is the same whether the decision is a metro system or a second retail unit. This section is about doing it when you are not a ministry.
The outside view earns its keep where three conditions hold together: the decision is consequential enough that being wrong hurts, the outcome depends on things nobody can know in advance, and the organisation has some history - its own or its industry’s - with decisions of that shape. That describes a great many ordinary commercial decisions: opening a second location, taking on debt, a first international expansion, a large systems implementation, a hiring plan built on a revenue ramp, a factory or fit-out project, a product launch into an adjacent category. Our use-case guides cover several of these individually, from opening a second location to taking on debt and investing in inventory.
Flyvbjerg’s own guidance on where the comparative advantage of the outside view is greatest is directly relevant to smaller organisations, and it cuts against intuition. The advantage is most pronounced for non-routine projects, understood as projects that managers and decision makers in a certain locale or organisation have never attempted before - “like building new plants or infrastructure or catering to new types of demand.” It is in the planning of such new efforts that the biases toward optimism and strategic misrepresentation are likely to be largest. He adds the reassuring observation that “most projects are both non-routine locally and use well-known technologies,” and are therefore particularly likely to benefit.
That is precisely the profile of the big decisions a growing business faces. A second location is novel to the company and utterly routine to the world. There are thousands of second locations in the record, and the fact that yours is your first is exactly what makes your internal estimate untrustworthy and the outside view valuable. The reflex - “we have never done this, so we have no data” - has the logic backwards. You have no *inside* data. The outside data is abundant, and its abundance is the whole point.
The practical starting point is an afternoon’s archaeology. Pull every significant commitment the organisation has made in the last five years for which there is both an approved plan and a known outcome: capital projects, system implementations, hires against a plan, launches against a forecast, budget versus actual on anything sizeable. For each one, record two numbers - what was approved, and what happened - plus the date and a one-line description. That is a reference class. It will be small, uneven and uncomfortable to read, and it will tell you more about your organisation’s estimating bias than any process improvement you could run.
Two features make this exercise disproportionately powerful compared to using industry data. The first is that it measures your competence, not an industry average, which means the correction it produces is genuinely yours and cannot be waved away as irrelevant. The second is political: it is very hard for a leadership team to argue that its own last eleven projects are not comparable to its next one. External benchmarks invite the response that the sample is not like us. Internal history does not.
Where internal history runs out, layer on the public categories. The Treasury table is generic by design and its categories - standard and non-standard buildings, standard and non-standard civil engineering, equipment and development, outsourcing - map onto commercial projects with very little strain. A software implementation is an equipment and development project. A warehouse fit-out on a complicated site is non-standard building work. Using the published upper bound as a starting point and adjusting downward only on evidence is a defensible method for an organisation with no data of its own, and it is exactly what the guidance intends when it refers to using “the best available data” in the absence of a specific evidence base.
With eight or twelve observations you cannot claim a P90 with a straight face, and pretending otherwise imports a false precision that undermines the credibility of the whole exercise the first time someone with a statistics background looks at it. Say what you have. A class of eleven projects supports a statement like: “our own history shows a median overrun of about a quarter, the worst was a bit over double, and none came in under budget.” That is a defensible, useful and honest summary, and it will change a funding decision.
Three disciplines make small classes more trustworthy. Report the raw observations alongside the summary, so readers can see what the percentile is made of. Use ranges rather than single percentiles - “between 45 and 70 per cent at the level of certainty we want” - because a small sample genuinely does not resolve finer than that. And bias your choice upward where the sample is thin, on the reasoning that small samples systematically miss the tail: with eleven observations from a right-skewed distribution, the chance that none of them is a tail event is high, so your observed maximum is probably not the real maximum.
The long-run payoff of running an outside view is that it makes your organisation’s estimating quality measurable, and therefore improvable. Every forecast you make with a stated confidence level becomes a testable claim, and once you have made enough of them you can ask the only question that matters about a forecasting process: when we said we were eighty per cent confident, how often were we right? Teams that track this typically discover that their eighty per cents come true perhaps half the time, and that discovery is worth more than any single corrected estimate.
This is also the only legitimate route to the discount discussed earlier. The method permits a lower uplift where there is evidence the team is significantly better than the class - and a calibration record is what that evidence looks like. Without one, a claim to be better than the class is indistinguishable from the optimism the method exists to correct. Incertive’s calibration tracking is built for exactly this loop: record the prediction with its confidence level, record what happened, and let the gap between the two become an organisational fact rather than a matter of opinion. Our guide to building a risk-aware culture covers the human side of making that habit stick.
Reference class forecasting is not difficult to understand, and that is misleading, because it is genuinely difficult to adopt. The obstacles are rarely technical. They are the predictable reactions of organisations to a method whose entire purpose is to take a number out of the hands of the people who wanted a different number. It is worth working through the objections honestly, including the ones that are partly right.
This is the universal first response, and it is not stupid - every project genuinely is different in some respects, and the method depends on comparability. But it is worth being precise about what the objection would have to establish to succeed. It is not enough that the project has unusual features. It must be shown that those features change the *distribution of outcomes*, and in a favourable direction, by more than the class variation already accounts for. Almost no version of this argument survives that test, because the class already contains projects that each had their own unusual features and still landed where they landed.
Kahneman’s curriculum story is the cleanest demonstration. The expert was asked directly whether the present team was better than the comparable ones, and said no - and the team proceeded on the inside-view estimate anyway. The uniqueness objection is rarely a claim about evidence; it is usually a claim about identity, an assertion that we are not the sort of organisation that those numbers describe. The productive response is not to argue but to ask what would have to be true for the class to be inapplicable, write it down, and check whether it is true. Occasionally it is, and then you have a better class. Usually the act of writing it down settles the matter.
This objection has real content. A class of forty-six rail projects genuinely does not know about your specific route, your contractor or your political environment, and a crude class applied mechanically can produce a number that is wrong for identifiable reasons. The mistake is treating this as an argument against the outside view rather than an argument for improving it. If the class is too crude, refine it - with statistical evidence that the refinement matters, as discussed above. If it cannot be refined for lack of data, the appropriate response is a wider stated range or a higher percentile, not a return to the inside view whose track record is the reason you are here.
It also helps to be clear about what standard the method has to meet. It does not have to be right. It has to be *less wrong, in a less predictable direction*, than the alternative. An estimate that is crude but unbiased beats one that is detailed and systematically thirty per cent low, and the record of the detailed alternative is not in doubt.
This is the strongest objection to the method, and Flyvbjerg raises it himself rather than waiting for critics to. He acknowledges that reference class forecasting may result in reserves large enough to create their own inefficiency, noting that for some projects the total budget reservation including uplifts “would be more than adequate,” and that “this may in itself create an incentive which works against firm cost control if the total budget reservation is perceived as being available to the project and its contractors.” Reserves will be spent simply because they are there, as the saying goes in construction.
The remedy is structural rather than analytical, and it is a distinction every organisation adopting the method needs to build in: the funded amount and the target are not the same number. Fund at the percentile the class requires; manage the project to a tighter internal target; hold the difference centrally, released only against demonstrated need. Flyvbjerg’s own prescription is to combine the uplifts with “tight contracts, maintained incentives for promoters to undertake good quantified risk assessment and exercise prudent cost control during project implementation.” The outside view sets how much you must be prepared to spend. It says nothing about how much you should aim to spend, and conflating the two turns a forecast into a licence.
The hardest barrier is not intellectual at all. Where inaccurate forecasts serve someone’s purpose, a better forecasting method is not a solution but a threat. Flyvbjerg’s analysis of this case is unsentimental: where strategic misrepresentation dominates, managers and forecasters “may not be interested” in accuracy “because inaccuracy is deliberate,” and “biased forecasts serve strategic purposes that dominate the commitment to accuracy and truth.” In that situation the potential for reference class forecasting is low and the barriers are high, because the demand for accuracy is not there.
His example is competitive and directly transferable to corporate life. Cities compete fiercely for scarce national funds, so pressures are strong to present projects as favourably as possible; there is no incentive for an individual city to debias, because unless all the others debias too, the honest one loses the competition. Substitute business units competing for capital allocation, or vendors bidding for work, and the structure is identical. Honest forecasting is individually costly and collectively beneficial, which is a coordination problem, not a methodology problem.
The answer, Flyvbjerg argues, is accountability rather than technique: measures that “reward accurate forecasts and punish inaccurate ones,” with forecasters and promoters made to carry the full risks of their forecasts, and independent bodies - national auditors or independent analysts - reviewing the work, which is itself a use for reference class forecasting. Inside a company the analogues are obvious and rarely implemented: score every business case against outcomes, publish the scores by sponsor, and make a sponsor’s forecasting record part of how their next proposal is read. An organisation that never looks back at its approvals is one where optimistic forecasting has no cost, and behaviour follows cost.
One genuine limitation deserves acknowledgement rather than defence. Flyvbjerg concedes that the outside view, “being based on historical precedent, may fail to predict extreme outcomes, that is, those that lie outside all historical precedents.” A class cannot price a category of event it has never seen. This is a real constraint, and it is why the method is a floor on prudence rather than a guarantee.
Two things blunt it in practice. First, as the Edinburgh report emphasises, most extreme projects are not caused by unprecedented catastrophes but by multiple ordinary adverse events occurring simultaneously - and those combinations *are* in the class, which is why the class must retain its outliers. Second, the response to genuine tail exposure is management rather than measurement: the report recommends early warning indicators, active attempts to break down project size, removing complexity, and preparing responses to and recovery from adverse events. Breaking a large irreversible commitment into a sequence of smaller reversible ones is the most reliable protection against a tail nobody can forecast, and it is available to organisations of any size.
Reference class forecasting was designed in an era when assembling a distribution meant commissioning a study, and where the output was a static table of uplifts published by a ministry. The logic is unchanged, but the mechanics have moved. What was once a multi-month consulting engagement is now a workflow that a competent team can run in an afternoon, and the outside view fits inside modern probabilistic tooling rather than sitting beside it as a separate governance ritual.
In a modern decision intelligence platform, a decision is represented as a model: a set of uncertain drivers, each with a distribution, combined by simulation into a distribution of outcomes. Reference class forecasting enters that model at two distinct points, and it is worth being clear about which you are doing.
The first is input calibration. Rather than asking the team for a range on each driver and accepting it, you set the ranges from class data wherever class data exists - historical variance in delivery duration for work of this type, historical variance between forecast and actual demand, historical spread on unit costs. The simulation then combines evidence-based ranges rather than judgement-based ones, and the output distribution inherits the honesty of its inputs.
The second is output correction. Where you cannot decompose the class into drivers - the usual situation, since published classes report total overrun rather than its components - you run the model on inside-view inputs and then compare its output distribution to the class distribution as a whole. If the model’s P80 sits well inside the class’s P50, the model is understating risk and the class is the authority. This is the check the Edinburgh case argues for most forcefully: not a replacement for the simulation, but an external validity test that the simulation cannot perform on itself.
Doing both, where you can, is better than either. Calibrated inputs make the model structurally sound; the output comparison catches the residual optimism that survives calibration, particularly the correlations and unlisted risks that no driver-level range captures. Our guide to sensitivity analysis covers the third leg - once the model is honest, finding out which drivers actually move the answer, so that mitigation effort goes where it changes the odds.
Most organisations attempt this in a spreadsheet, and most give up. The obstacles are practical rather than conceptual: spreadsheets model point values natively and distributions only with effort, correlations between drivers are painful to express, re-running a scenario means rebuilding rather than re-simulating, and the reference class data - the thing that most needs to persist and grow - ends up in a tab that nobody maintains after the deal closes. Our analysis of the limits of spreadsheet forecasting covers the failure modes in detail.
The deeper problem is institutional. A reference class is an asset that appreciates: every completed project adds an observation, tightens the distribution and strengthens the organisation’s claim to know its own performance. That only happens if the class lives somewhere durable, is updated as a matter of routine, and is consulted by default at approval. A file on someone’s drive satisfies none of those conditions, which is why organisations that do this well treat the class as part of their planning infrastructure rather than as an analysis someone once did.
However the analysis is produced, the presentation determines whether it changes anything. A board does not need the distribution; it needs the four things the distribution implies. What is the probability we hit the target as currently planned? What would we have to fund to be as certain as we want to be? Which drivers are responsible for most of the spread? And what specifically would move the odds?
Those four questions map onto four outputs: a success probability, a percentile-based funding figure, a tornado diagram ranking the drivers, and a set of recommended changes. Delivered in that form, an outside-view analysis stops being an argument about pessimism and becomes a set of decisions: how certain do we want to be, what does that cost, and where do we spend effort to make it cheaper. Incertive is built to produce exactly that from a plain-language description of the decision, and you can see the shape of the output in a sample analysis before modelling anything of your own.
The gap between understanding reference class forecasting and practising it is mostly organisational. What follows is a sequence that works in most companies, ordered so that each step produces something useful on its own - because a programme that requires eighteen months of data collection before it delivers value will be cancelled at month four.
One caution about sequencing. Resist the urge to begin with a policy. Reference class forecasting introduced as a mandate arrives as a tax on business cases and is resented accordingly; introduced as an analysis that changed one important decision for the better, it tends to be requested rather than imposed. The Treasury could mandate it because the Treasury controls the money. Most people reading this cannot, and do not need to.
If you want structure to start from, our risk planning template and project success calculator are built around the same logic, and the methodology page sets out how Incertive combines outside-view calibration with simulation in a single workflow.
The argument for reference class forecasting comes down to a single asymmetry. Inside-view estimating has been measured, repeatedly, across seventy years, tens of thousands of projects and every sector anyone has bothered to study, and it is wrong in a consistent direction by margins that routinely change whether a project should have been approved. The outside view has been tested against it and wins - not because it is clever, but because it declines to rely on the faculty that keeps failing.
What the method actually delivers is narrower than its advocates sometimes suggest and more valuable than its critics allow. It will not tell you what will go wrong on your project. It will not price a risk that has never occurred anywhere. It will not fix an organisation whose incentives reward optimistic numbers, though it will make the optimism visible, which is a start. What it will do is replace a number somebody chose with a number derived from what happened to everyone else who tried this, and attach to it a stated level of confidence that someone must consciously accept.
That is a modest-sounding change with immodest consequences. It moves contingency from convention to evidence. It makes the cost of certainty explicit, so that risk appetite becomes a decision rather than an accident. It gives the person who suspects an estimate is too low something better than a suspicion. And over time, as the organisation scores its own outcomes and builds its own classes, it converts forecasting from a matter of temperament into a measurable capability that can actually improve - which no amount of exhortation to be realistic has ever achieved.
The Edinburgh Tram remains the cleanest illustration of what is at stake. In 2004 the promoter’s number and the reference class forecast were both on the record, roughly eighty million pounds apart, and the outside view came with a caveat that it was probably still too low. The project proceeded on the inside view. The eventual real-terms overrun was 52%, on a line substantially shorter than the one approved. The information required to know better was available, published, and ignored - which is the ordinary fate of an outside view that arrives after everyone has decided what they want the answer to be. The remedy is to ask the question before that point, every time, as a matter of routine: what happened to everyone else who did this?
Reference class forecasting is a method for producing an unbiased forecast by taking the “outside view”: instead of estimating a project from its own details, you identify a reference class of comparable completed projects, establish the distribution of their actual outcomes against their approved budgets and schedules, and position your project within that distribution. Bent Flyvbjerg set out the method in three steps in the Project Management Journal in 2006, building on Daniel Kahneman and Amos Tversky’s work on decision-making under uncertainty. Because it reads outcomes directly off history rather than reasoning forward from assumptions, it bypasses optimism bias and strategic misrepresentation rather than trying to argue an estimator out of them.
A risk register and a conventional Monte Carlo simulation are both inside-view methods: they quantify the risks a team was able to think of, using ranges the same team supplied. They are valuable, but they inherit whatever the team failed to imagine, and their outputs are only as unbiased as the judgements behind them. Reference class forecasting works from recorded outcomes instead, so it captures the causes of overrun nobody listed - including the ones that were unknowable in advance. The two are complements: the expert report to the Edinburgh Tram Inquiry concluded that an inside-view risk register plus Monte Carlo simulation was insufficient on its own and could create a false sense of certainty, and recommended the outside view as the check on it.
There is no universal minimum, and the honest answer is that it depends on how much spread the class has: the wider the distribution, the more observations you need before the percentiles are stable. The published guidance offers a sense of scale - the UK Department for Transport guidance was built from a sample of 260 transport projects, including a reference class of 46 rail projects, and the road distribution reported by Flyvbjerg drew on 172 completed projects. In a private company you will rarely have anything like that, and a class of ten or fifteen genuinely comparable past projects, honestly measured, still beats an inside-view estimate with no outside check at all. What matters more than raw count is that the comparison is like-for-like and that no project was quietly excluded for being embarrassing.
An optimism bias uplift is the percentage you add to an estimate to correct for the systematic tendency of appraisers to be too optimistic - the practical output of a reference class forecast. HM Treasury’s supplementary Green Book guidance publishes indicative upper-bound capital-expenditure uplifts by project type, including 24% for standard buildings, 44% for standard civil engineering, 51% for non-standard buildings, 66% for non-standard civil engineering and 200% for equipment and development projects, with appraisers instructed to start at the upper bound and reduce it only where the contributory factors have demonstrably been mitigated. The size of the uplift depends on the level of certainty you want: the UK Department for Transport uplifts for road schemes rise from 15% at P50 to 32% at P80 and 45% at P90.
Usually yes, because uniqueness is far rarer than it feels. Flyvbjerg’s argument is that the outside view is most valuable precisely for non-routine projects - the ones an organisation has never attempted before - because that is where optimism bias and strategic misrepresentation are largest. What matters is not that your project is identical to the class but that it is statistically similar in the drivers that determine overrun: scale, novelty, interface complexity, procurement route, stakeholder count. Genuine novelty makes choosing the class harder, and it does mean the outside view may fail to predict outcomes outside all historical precedent; it does not make the inside view more reliable. The practical response is a wider class and a higher percentile, not a return to guessing.
Incertive turns a plain-language description of a decision into a probability of success, the drivers responsible for the risk, and the changes that most improve your odds - so contingency stops being a round number and becomes a choice you can defend.
Analyze My DecisionBack to Blog