Capital project risk is the gap between what a project is approved to cost and what it will actually cost. This guide turns that gap into a number: the base rates, the drivers, and the contingency you can defend.
Capital project risk is the exposure that a large, largely irreversible investment will cost more, take longer, or return less than the case on which it was approved. It is the most consequential risk most organisations carry, because a capital project concentrates years of cash flow into a single commitment that cannot be unwound cheaply once concrete is poured or steel is ordered. And it is the risk most often managed with the weakest tools: a register of hazards, a colour-coded matrix, and a contingency percentage that nobody in the room can actually justify.
This guide treats capital project risk as what it is - a measurable quantity, not a mood. It sets out the base rates that the evidence supports, where the exposure actually comes from, why a register and a heat map cannot size it, how to quantify it with ranges and simulation, how to set a contingency you can defend to a board, how to rank the drivers so mitigation money goes where it works, and how to carry the whole discipline through the stage gates from screening to sanction and into execution. Where a figure appears, it is either linked to the source we verified or explicitly labelled illustrative.
This post sits inside a wider cluster on managing delivery risk. If you are new to the subject and want the foundations first, start with our pillar guide to project risk management for non-technical leaders, then come back here for the capital-project-specific treatment: the sanction decision, the front-end loading problem, contingency at a chosen confidence level, and the sector base rates that should anchor every estimate you approve.
The phrase gets used loosely, which is part of the problem. In many organisations “capital project risk” means the document that lists things that might go wrong: a spreadsheet of hazards, each with an owner, a likelihood rating, an impact rating, and a mitigation sentence. That document has its uses, and we will come to them, but it is not the risk. The risk is the gap between what a project is approved to cost and take, and what it will actually cost and take - a gap that exists whether or not anyone writes it down, and that has a size and a shape that can be estimated with the same rigour as the base estimate itself.
Put formally, capital project risk is the probability distribution of outturn around the sanction case. The sanction case is a set of promises: a total installed cost, a completion date, a production or service capability, and a stream of benefits that together justify the investment. Each of those promises rests on hundreds of assumptions, and every assumption has a range. Risk is what happens when those ranges combine. It is not a synonym for “bad things”; a project can come in under budget, and the same analysis that tells you the chance of an overrun tells you the chance of an underrun. What makes capital projects distinctive is that the distribution is usually skewed: the upside is bounded by physics and contract, while the downside has a long tail.
The most useful mental shift is from enumeration to measurement. A hazard list answers the question “what could go wrong?” Exposure answers the question “how wrong could it go, and how likely is that?” These are different questions and they require different machinery. You can have an immaculate register with fifty well-owned risks and still have no idea whether your budget is adequate, because the register never adds up to a number. Conversely, you can have a rough register and a rigorous quantification and know, with defensible precision, that you have a thirty percent chance of exceeding the approved cost by more than fifteen percent - which is the sort of statement that changes what a board decides.
Exposure also includes the variability that never makes it onto a register at all. Registers capture discrete, nameable events: a permit is refused, a contractor fails, a key piece of equipment arrives late. But most cost growth on capital projects does not come from named events. It comes from the accumulation of ordinary estimating error across thousands of lines - quantities that were slightly low, rates that drifted, productivity that ran below the assumed norm, small scope additions that each seemed reasonable. None of these is a “risk” in the register sense. All of them are exposure. Any method that only measures the named events will systematically understate the real range, which is one reason organisations are so often surprised by overruns they had, technically, documented nothing about.
Framing risk as exposure also clarifies who owns it. A hazard has an owner because someone must act on it. Exposure is owned by the sponsor and ultimately the board, because it is the thing being underwritten when capital is released. When a project is approved at a number, the organisation is implicitly accepting the whole distribution around that number, including the tail. Making the distribution explicit does not create the exposure; it simply stops the organisation from accepting it blind. That is the entire argument for quantification, and it is worth stating plainly before we get to method.
Capital project risk is usually discussed as cost risk, because cost is the dimension that shows up in the accounts. But an investment case rests on four legs, and a failure in any of them destroys value. Cost risk is the chance the installed cost exceeds the sanctioned figure. Schedule risk is the chance the asset is not available when promised, which matters because a late asset is a late revenue stream and often a late set of contractual obligations. Scope and performance risk is the chance the asset that gets built is not the asset that was approved: a plant that runs at ninety percent of nameplate, a building that meets code but not the operating brief, a system that ships without the modules that carried the business case.
Benefits risk is the least measured and often the largest. A capital project is justified by what it will produce, and the forecast of that production is at least as uncertain as the forecast of its cost, usually more so, because it depends on demand and prices years out rather than on quantities and rates that engineers can bound. The research on this is uncomfortable. Analysis of a large sample of rail megaprojects found not only substantial cost overruns but demand forecasts that were badly optimistic, with actual passenger traffic falling far short of what had been promised at approval.
44.7% and 51.4% - across a sample of rail megaprojects, average cost overrun ran to roughly 44.7 percent while passenger demand was overestimated by about 51.4 percent - meaning the case was wrong at both ends at once.
Source: McKinsey, Megaprojects: The good, the bad, and the better
The four dimensions interact, which is why treating them separately understates the total. A schedule slip pushes cost up through prolongation and escalation, and pushes benefits back by delaying the revenue start. A scope reduction taken to protect cost can quietly remove the capability the benefits case depended on, converting a cost success into a value failure. The disciplines that measure only one leg - a cost risk analysis with no schedule model, or a schedule risk analysis with no cost consequence - miss the compounding. A serious treatment of capital project risk models the legs together, or at minimum makes the linkage between them explicit.
Operating decisions are frequent, small relative to the balance sheet, and reversible. If a pricing change underperforms you change it back next quarter; if a hire does not work out you have absorbed a few months of salary. Capital projects invert every one of those properties. They are infrequent, so the organisation has little practice and few internal base rates. They are large relative to the balance sheet, so a single bad outcome can consume years of earnings. And they are close to irreversible: once a foundation is poured, a long-lead order is placed, or a site is committed, the option to stop cheaply has expired.
That irreversibility is what makes the front end so valuable and so frequently underfunded. The period before sanction is the only window in which the cost of changing your mind is measured in engineering hours rather than in construction claims. It is also, perversely, the period in which the organisation knows least. Every dollar of front-end definition buys a narrower range at sanction, and a narrower range at sanction is worth far more than the same dollar spent optimising execution later. The industry term for this is front-end loading, and the evidence that skimping on it drives failure is among the more consistent findings in the field.
65% - megaprojects executed around the world have been found to carry roughly a 65 percent failure rate, with basic front-end loading errors - starting execution from poor levels of project definition - identified as a leading cause.
The frequency point deserves its own emphasis, because it explains why capital project risk resists the usual learning loop. An organisation that runs a thousand small experiments a year gets calibrated by feedback: it learns what its estimates are worth. An organisation that sanctions three major projects a decade gets almost no feedback at all, and what feedback it does get arrives years later, filtered through staff turnover and reorganisation. This is exactly the condition under which reference class forecasting earns its keep, because it substitutes the outside world’s data for the internal experience you do not have.
A distinction economists draw is between risk, where the possible outcomes and their probabilities are known, and uncertainty, where they are not. In practice capital projects sit in between, and the honest position is that you can usually bound the range even when you cannot pin the probability. You do not know what steel will cost in three years, but you can say with confidence that it will not be half or triple today’s price, and you can put a shape on the plausible middle. That is enough to compute with. The perfectionist objection - that we cannot know the true distribution, so we should not model one - has the practical effect of defaulting to a single number, which is a distribution too: the worst one, a spike with no width at all.
It is also worth separating aleatory variability, which is genuine randomness that more study will not remove (weather windows, ground conditions, market prices), from epistemic uncertainty, which is ignorance that study can reduce (an unsurveyed site, an undefined interface, an unsigned contract). The distinction is practical rather than academic: epistemic uncertainty is an argument for spending money at the front end to narrow the range, while aleatory variability is an argument for carrying contingency and for contractual mechanisms that allocate the variation to whoever is best placed to bear it. Confusing the two produces both wasted study and unfunded exposure.
Finally, some of what is called uncertainty is really undisclosed information. Estimates are produced inside organisations with incentives, and a number that has to clear a hurdle rate tends to clear it. This is not fraud in most cases; it is the ordinary pressure of wanting a project to proceed, applied to hundreds of small judgement calls that each lean the same way. We treat this properly in the section on behavioural drivers below, and at length in our guide to optimism bias in business, but it belongs in the definition: capital project risk includes the risk that the estimate itself is biased, not merely imprecise.
Before modelling anything, it is worth establishing what happens to capital projects in general, because that is the prior against which any specific estimate should be judged. If your project team hands you a number and the historical record says projects like yours land forty percent above their sanction estimate, the burden of proof sits with the team to explain what is different, not with the sceptic to explain why history might repeat. The base rate is the cheapest risk analysis available and most organisations never look it up.
The single most cited finding in the field is that cost overrun on large projects is not an occasional accident but the normal case. Work drawing on decades of infrastructure data has been summarised as an “iron law” of megaprojects: over budget, over time, over and over again, with roughly nine in ten projects experiencing a cost overrun.
9 in 10 - roughly nine out of ten megaprojects experience cost overrun, a pattern so consistent across sectors, countries and decades that it has been described as the Iron Law of Megaproject Management.
The finding is not an artefact of one dataset. An earlier and much-replicated study of 258 transport infrastructure projects across twenty countries found average cost escalation of around 28 percent, with clear differences by asset class: roads around 20 percent, bridges and tunnels around 34 percent, and rail around 45 percent. The same study made a point that matters more than the headline: the accuracy of forecasts had not improved over the seventy years for which data were available. This is the crucial claim, because it removes the easiest excuse. If overruns were a legacy of primitive tools, they would be shrinking. They are not.
~28% average - across 258 transport infrastructure projects in 20 nations, average cost escalation was about 28 percent, with roads around 20 percent, bridges and tunnels around 34 percent, and rail around 45 percent - and no evidence that forecasting accuracy had improved over seventy years.
Source: Flyvbjerg, Holm and Buhl, Underestimating Costs in Public Works Projects
The hydropower record makes the same point with even sharper numbers. A study of 245 large dams built across 65 countries found that construction costs ran on average about 90 percent above the budget approved at the decision to build, in real terms, and that mega-dams took an average of 8.2 years to construct, frequently more than ten. The authors were explicit that the magnitude of overrun had not declined over time: dam budgets today are, as they put it, as wrong as at any point in the seventy years of record.
+90% on 245 dams - across 245 large dams in 65 countries, actual construction costs averaged about 90 percent above the budget at the time of approval in real terms, with an average build time of 8.2 years and no decline in overrun magnitude over seven decades.
Source: Blavatnik School of Government, University of Oxford
Base rates differ by asset class, and using the right one matters. Building a reference class is the subject of its own guide, but a starting map of the published evidence is useful, and each of the following has been verified against its source. In construction, McKinsey’s work on large projects reports that they typically run up to 80 percent over budget and around 20 percent longer than scheduled, while a separate McKinsey analysis of more than 500 major projects put average cost overrun near 79 percent. In power and utilities, an EY study of 100 of the world’s largest generation, transmission, distribution and water projects found 64 percent delayed and 57 percent over budget, with the average project running 35 percent (about two billion US dollars) over and two years late.
57% over budget, 64% late - across 100 of the world’s largest power generation, transmission, distribution and water projects, 57 percent were over budget and 64 percent were delayed, with the average project running 35 percent (about US$2bn) over budget and two years behind schedule.
In building construction more broadly, KPMG’s global construction survey found that only 31 percent of respondents’ projects had come within ten percent of budget in the preceding three years, a figure that says as much about the width of the distribution as about its centre. In technology-heavy capital programmes, the picture is different in shape rather than in severity: research on around 1,500 IT projects found an average cost overrun near 27 percent, but with roughly one in six turning into a “black swan” at about 200 percent cost overrun and 70 percent schedule overrun. That fat tail is the defining feature of technology risk, and it is why an average is the wrong summary statistic for a technology-heavy capital programme.
1 in 6 is a black swan - across roughly 1,500 IT projects, the average cost overrun was about 27 percent, but one in six projects became a “black swan” with around 200 percent cost overrun and 70 percent schedule overrun - the average badly understates the tail.
Public-sector portfolios show the same signature at national scale. The United States Government Accountability Office publishes an annual assessment of NASA’s major projects, which is unusually useful because it is an audited, consistent, multi-year series rather than a survey. Its July 2026 edition reported a portfolio of 36 major projects, 18 of them in development, with an estimated life-cycle cost of at least 70 billion US dollars, carrying cumulative cost growth of 4.7 billion dollars and cumulative schedule delays of 14 years - and it noted that a single programme, the Orion crew capsule, accounted for almost three quarters of that cumulative cost growth.
$4.7bn and 14 years - NASA’s portfolio of 36 major projects carried cumulative cost growth of $4.7 billion and cumulative schedule delays of 14 years as of the July 2026 assessment, with the Orion crew capsule alone accounting for almost 75 percent of the cumulative cost overruns.
Source: U.S. Government Accountability Office, GAO-26-108556
A base rate is not a prediction of your project. It is a prior: the starting point you should hold before you look at project-specific detail, and the thing your project-specific detail has to argue you away from. Used properly it does two jobs. It calibrates the centre of your estimate, by telling you what projects of this type actually cost relative to what they were approved at. And it calibrates the width, by telling you how much spread to expect - which is usually far more than an internal estimating team, working bottom-up from a fully specified takeoff, will naturally produce.
The common objection is that every project is unique. In one sense that is true and in the sense that matters it is not. The concrete, the site, the client and the contract are unique; the human and organisational processes that generate estimates and then overrun them are not. It is precisely because the failure modes are generic that the base rates are so stable across sectors and decades. Teams that insist their project is exempt are usually invoking the inside view - the vivid, detailed, plan-shaped story of how this particular job will go - which is exactly the mode of thinking that the historical record says is unreliable. Our guide to reference class forecasting works through how to build the outside view in practice, including when no public database of your asset class exists.
Base rates are also not an excuse for fatalism. The purpose of knowing that nine in ten projects overrun is not to shrug at the tenth; it is to fund and govern the project so that it can be the tenth, and to be honest at sanction about the odds you are taking. The organisations that beat the base rate do a small number of unglamorous things consistently: they invest in front-end definition, they estimate as ranges, they set contingency at a chosen confidence level, they re-test at every gate, and they keep the outturn data so that the next estimate is better calibrated than the last. Every one of those is discussed below.
It is worth pausing on the size of the pool this applies to. Construction alone is one of the largest capital-forming activities in any developed economy. In the United States, the Census Bureau put total construction spending in June 2026 at a seasonally adjusted annual rate of about 2,166.5 billion dollars, of which roughly 745.3 billion was private nonresidential and 544.1 billion was public construction. Apply even a conservative reading of the overrun literature to a base of that magnitude and the aggregate value at stake in getting capital project risk right is measured in hundreds of billions annually.
$2.17 trillion - total US construction spending ran at a seasonally adjusted annual rate of about $2,166.5 billion in June 2026, including roughly $745.3 billion of private nonresidential and $544.1 billion of public construction.
The organisational version of the same point is more immediate. For most companies, the capital programme is where the balance sheet is actually deployed, and the difference between a portfolio delivered at plan and one delivered at plan plus thirty percent is the difference between funding the next round of investment from operations and going back to the market. That is why capital project risk deserves board-level quantification rather than a schedule appendix, and why the rest of this guide is concerned with turning it into numbers a board can act on. For the broader argument about what poor decision quality costs an organisation, see the hidden costs of bad business decisions.
Overruns are usually explained after the fact by a single villain: a difficult contractor, a hostile regulator, an unforeseeable commodity spike. The record does not support such tidy accounts. Cost and schedule growth on capital projects is nearly always the accumulation of many partly independent, partly correlated pressures, most of which were visible in outline at sanction and none of which was individually large enough to trigger alarm. Naming the sources properly is the first step to modelling them, because you cannot put a range on a driver you have not identified.
The dominant source of capital project risk is the immaturity of the estimate on which the project is approved. Estimating practice has long recognised this by classifying estimates according to how well the scope is defined, from a screening-level figure produced from capacity factors and analogies through to a control estimate built from a near-complete design with quantified takeoffs and firm quotations. The classes exist precisely because the expected accuracy of an estimate is a function of definition maturity, and the accuracy band around a screening estimate is wide enough to swallow most investment cases whole.
The practical failure is not that organisations produce early-class estimates; they must, because you cannot fully define a project you have not yet decided to do. The failure is that the class is forgotten by the time the number reaches the approval paper. An estimate produced from a five percent design, appropriately caveated by the estimator with a wide band, is copied into a business case as a single figure, discussed at a committee as a single figure, and approved as a single figure. The band evaporates somewhere between the estimating department and the boardroom, and with it the only honest signal about how much the organisation is really committing to. This is the hidden cost of false precision in its purest form.
The remedy is procedural before it is analytical: require that any estimate presented for a decision carries its class and its band, and require that the band be quantified rather than asserted. If a project is being sanctioned at a stage where the honest range is minus twenty to plus fifty percent, the board should be told that, because the appropriate governance response - a smaller commitment now, a staged release, more front-end money, or a decision to wait - depends entirely on knowing it. Approving a wide-band estimate is a legitimate choice; approving it while believing it is narrow is not a choice at all.
Closely related, but distinct, is the completeness of the scope itself. An estimate can be well built and still be an estimate of the wrong project, because the project as defined at sanction omits work that will unavoidably be required: enabling works, interfaces with existing operations, temporary facilities, commissioning spares, the integration of a system with the six other systems it must talk to. Scope omission is not the same as scope creep. Creep is discretionary addition after sanction; omission is work that was always necessary and was never counted. Omission is more dangerous because nobody can be refused for asking for it.
Front-end loading is the disciplined counter to both. It means completing a defined set of definition activities before sanction: site investigation, basic engineering, interface identification, constructability review, procurement strategy, operating philosophy, and a genuine execution plan. The evidence that skipping it drives failure is strong, as the megaproject failure findings cited earlier make clear, and the mechanism is intuitive. Every undefined element at sanction is a placeholder, and placeholders in capital projects almost always resolve upward, because the pressure at the moment of resolution is to keep the work moving rather than to re-open the case.
The organisational obstacle is that front-end money is spent before there is a project to charge it to, and is therefore the easiest budget to cut. A useful reframing for a sponsor is that front-end spend is not a cost of the project; it is the purchase of a narrower distribution. Quantifying that trade directly - modelling the outturn range with and without a completed site investigation, for example - turns an argument about whether to spend two million dollars into an argument about whether two million dollars is worth a materially tighter band on a four hundred million dollar commitment. Framed that way it usually answers itself.
Capital projects buy commodities, fabricated equipment, and labour hours in markets that move, over durations long enough for the movement to matter. Escalation is often handled with a single index applied to the whole estimate, which conceals two problems. The first is that escalation is not uniform: structural steel, copper, transformers, specialist labour and shipping have their own cycles and their own lead-time behaviour, and the components with the longest lead times are frequently the ones with the most volatile prices. The second is that escalation interacts with schedule: a project that runs late escalates for longer, so schedule risk and cost risk are correlated through the market as well as through prolongation.
Supply chain exposure has become more prominent as capital programmes concentrate on a narrow band of specialised equipment. Where a single class of long-lead item has a small number of qualified suppliers globally, the effective risk is not a price range but a queue position, and the appropriate model is closer to a delay distribution with a fat right tail than to a symmetric price band. The practical implication for quantification is that these items deserve their own treatment rather than being buried in a bulk allowance, because their variance dominates and their behaviour is not well described by the same distribution as the rest of the estimate.
Contracting strategy allocates market exposure but does not eliminate it. A lump-sum turnkey contract transfers price variability to a contractor, who prices the transfer, and who may not survive the outcome if the transfer proves badly mispriced. Contractor insolvency is a real and periodically severe source of capital project risk, and it is precisely correlated with the market conditions that make the transfer valuable. Any model that treats a fixed price as risk-free is missing the counterparty’s own distribution, which is a version of the correlation problem discussed below.
Once work starts, the largest single lever on cost is usually productivity: how many units of output a crew actually achieves per hour compared with the norm assumed in the estimate. Productivity is sensitive to congestion, sequencing, weather, rework, information availability and morale, and it varies far more than most estimates allow. Because labour hours multiply through a large portion of the direct cost, a modest productivity shortfall converts into a substantial cost movement, and it does so quietly, showing up as a gradual erosion of progress against plan rather than as a discrete event that anyone reports as a risk.
Interfaces are the second execution multiplier. Every boundary between packages, contractors, disciplines or systems is a place where information must transfer correctly and on time, and the number of such boundaries grows faster than the number of packages. Interface failures generate rework, and rework is doubly expensive because it consumes the same constrained resource twice while pushing everything downstream. Schedule models that treat activities as independent miss this entirely, which is one of the reasons schedule risk analyses so often produce optimistic answers even when the individual duration ranges are honest.
Execution risk is also where the schedule and the cost model must be joined. Delay costs money through prolongation of site overheads, extended equipment hire, financing during construction, and escalation, and it delays benefits. Modelling the schedule and cost as separate exercises with separate contingencies double-counts some exposure and misses more of it. Our guide to Monte Carlo simulation in project management works through how the two models are linked in practice, including the treatment of parallel paths and the merge bias that makes optimistic completion dates so persistent.
For many capital projects the critical path runs through a consent, not a construction activity. Permitting risk has an awkward statistical shape: much of the time the process completes within a normal band, and occasionally it is extended by a challenge, a judicial review, an election, or a change in policy, producing delays measured in years rather than months. A symmetric range around a planning assumption models this badly. The right representation is usually skewed, with a long right tail and a floor set by the statutory minimum, and it is one of the few drivers where a discrete-event treatment sits more naturally than a continuous range.
Stakeholder risk is adjacent and often underestimated because it is uncomfortable to write down. Communities, adjacent landowners, environmental groups, unions and political sponsors all have the capacity to change a project’s cost and schedule, and their behaviour is partly a function of how the project engages them. Unlike commodity prices, this is a driver that project decisions materially influence, which makes it a prime candidate for the kind of mitigation testing described later: model the outturn with and without an early engagement programme and see what the probability of on-time consent does.
Regulatory change during a long build is the third strand. Codes, emissions standards, safety requirements and tax treatments move, and a project sanctioned under one regime may be commissioned under another. This is genuine aleatory uncertainty from the project’s point of view and belongs in the model as such. It is also a reason to prefer designs with headroom against foreseeable tightening, a choice that costs money at sanction and buys optionality later - exactly the sort of trade that a probabilistic model can price and a single-point estimate cannot.
The dimension least often quantified is the one that determines whether the asset was worth building. Benefits depend on demand, price, availability and operating cost, all forecast over an asset life that may run for decades. The rail evidence cited earlier - demand overestimated by roughly half - is not an outlier peculiar to railways; it is what happens when a forecast is produced by a party that wants the project approved and is never subsequently audited against outturn. Modelling benefits as a range, and reporting the probability that the investment clears its hurdle rate rather than a single net present value, changes what gets approved.
Benefits risk also interacts with the decision structure. Many capital projects are effectively bets on a demand scenario, and the appropriate response to demand uncertainty is often not more contingency but a different design: modularity, phasing, or the deliberate purchase of expansion options. A probabilistic model makes those alternatives comparable, because a phased scheme with lower expected value and much lower variance can be shown to have a higher probability of clearing the hurdle than a monolithic scheme with a better central case. That comparison is invisible in a single-point net present value, which is why so many organisations default to the monolith.
For decisions where the benefits case is the crux rather than the construction cost, the analysis is closer to what we describe in scenario analysis and probabilistic forecasting: the machinery is the same, but the uncertain drivers are commercial rather than physical. The best capital governance runs both views on the same project, so that a scheme with a robust construction estimate and a fragile demand case is not mistaken for a low-risk investment.
Underneath every technical driver sits a human one. Estimates are made by people who want their project to happen, reviewed by people who want the pipeline to look healthy, and approved by people who have already been told the project is a good idea. The result is a systematic downward pressure on cost forecasts and upward pressure on benefit forecasts that no amount of bottom-up rigour corrects, because the bias operates on the hundreds of judgement calls that bottom-up rigour is made of.
This is well enough established that it has been written into public appraisal rules. HM Treasury’s supplementary Green Book guidance states plainly that project appraisers tend to be over-optimistic and that explicit adjustments should be made to estimates of a project’s costs, benefits and duration, providing cost and time uplift percentages for generic project categories to be used in the absence of more robust primary evidence. It is worth dwelling on what that means: a national treasury has concluded that the correct default is to assume the estimate in front of you is too low and to add to it as a matter of policy.
Policy, not opinion - HM Treasury’s supplementary Green Book guidance states that project appraisers have a tendency to be over-optimistic and that explicit uplifts should be applied to estimates of costs, benefits and duration in the absence of robust primary evidence.
Source: HM Treasury, Green Book supplementary guidance: optimism bias
The distinction between innocent optimism and strategic misrepresentation matters less than the correction, which is the same in both cases: anchor on the outside view, quantify the range, and make the person presenting the number state the probability that it is sufficient. Asking “what is the chance this budget holds?” is a more productive governance question than “are you confident?”, because the first has an answer that can be wrong in a checkable way. Our guides to optimism bias in business and why business plans fail go further into the mechanisms; for capital projects, the operative point is that behavioural risk is a driver like any other and belongs in the model rather than in the postmortem.
Nearly every capital project maintains a risk register, and most maintain a probability-and-impact matrix that colours those risks red, amber and green. These artefacts are so universal that questioning them feels heretical, so it is worth being precise: the register is a good tool doing a job it was designed for, and a bad tool for the job it is usually asked to do. It manages attention. It does not measure exposure. Confusing the two is the single most common analytical failure in capital project governance.
A register is a management instrument. It records that someone has thought about a threat, names an owner, records an action and a date, and provides a forum in which those actions get chased. That is valuable and no quantitative method replaces it. Risks that are being actively worked need a place to live, and a project without a register tends to be a project where accountability for threats is diffuse. The register is also a decent memory: it captures the reasoning of a team at a point in time and lets a later reviewer see what was anticipated and what was not.
Where registers earn their reputation for uselessness is when they become compliance artefacts: reviewed monthly, scored by habit, with the same forty entries carrying the same amber ratings for two years while the project quietly drifts. The failure mode is easy to spot. If no entry has ever been closed because the threat has passed, and no entry has ever been escalated because the score moved, the register is documentation rather than management. That is a governance problem, not an argument against registers, but it does explain why so many boards have stopped reading them.
The deeper problem is arithmetic. A typical matrix scores probability on a one-to-five scale and impact on a one-to-five scale and multiplies them to get a score between one and twenty-five. These scales are ordinal: a “4” for probability is more likely than a “3”, but it is not a number in any sense that supports multiplication. Multiplying two ordinal labels produces a value with no units and no meaning, and the resulting ranking is an artefact of how the bands were drawn rather than a property of the risks. Change the band boundaries and the ranking changes, though nothing about the project has.
Because the scores have no units, they cannot be aggregated. There is no valid way to combine a portfolio of matrix scores into a statement about the project as a whole, which means the register cannot answer the only question the board is really asking: is the budget enough? Organisations paper over this by summing scores or counting reds, both of which are meaningless operations, and then by setting contingency at a percentage that has no relationship to either. The chain from analysis to decision is broken at the first link and everything downstream inherits the break.
Quantitative analysis fixes this by insisting that risk be expressed in the units of the thing at stake. A driver is not a “high” risk; it is a range of dollars or a range of days. Ranges in real units can be combined, because money adds and durations sequence. Once every driver is expressed that way, the aggregate is computable and the question of budget adequacy has an answer. This is the essential difference between a qualitative and a quantitative treatment, and it is why our guide to risk analysis software frames the maturity ladder as a progression from listing to scoring to measuring.
Even teams that do quantify often quantify wrongly, in one of two symmetrical ways. The first is to add worst cases: take the pessimistic value of every line and sum them. This produces a figure so large that no project could ever be approved against it, and it is also statistically wrong, because the probability that every driver simultaneously lands at its worst is vanishingly small unless the drivers are perfectly correlated. Teams that do this once usually never do it again, and the overcorrection is the second error: add the most likely values instead, which understates the total, because the sum of most-likely values is not the most likely sum when the distributions are skewed.
The correct treatment is to sample. Draw a value for each driver from its own distribution, compute the total for that draw, and repeat many thousands of times. The resulting distribution of totals reflects both the individual uncertainties and the fact that they will not all land badly at once. This is what Monte Carlo simulation does, and it is the reason it produces answers that are neither the paralysing worst case nor the falsely comfortable central case. Our three-point estimation guide covers the mechanics of getting from a low, likely and high estimate to a distribution that can be sampled.
Correlation is where naive sampling goes wrong and where capital projects differ most from simpler models. Drivers on a construction project are rarely independent. If labour productivity is poor, it is probably poor across several packages at once, because the causes - congestion, weather, a shortage of skilled trades in the region - are shared. If steel prices are high, they are probably high for all steel packages simultaneously. Independent sampling of correlated drivers cancels the variability out and produces a distribution that is far too narrow, which is worse than no analysis at all because it carries the authority of a model. Any serious treatment of capital project risk has to represent the drivers that move together, and any tool that cannot represent correlation should be treated with suspicion.
The end point of a qualitative process is almost always a percentage: add ten percent for contingency, or fifteen, or whatever the organisation’s convention happens to be. The percentage feels prudent, and it has the enormous political advantage of being uncontroversial, because it is the same number as last time. But it encodes no information about the project it is applied to. The same fifteen percent covers a well-defined repeat build on a known site and a first-of-a-kind facility with an unsurveyed foundation, which cannot both be right.
A percentage contingency also gives no answer to the question that determines whether it is adequate: what probability of sufficiency does it buy? Without a distribution, nobody can say whether the fifteen percent covers the project at a fifty percent confidence level or an eighty-five percent one, and so nobody can say whether the organisation is taking a prudent or a reckless position. Two projects in the same portfolio carrying the same percentage may be at wildly different confidence levels, which makes portfolio-level risk management impossible even in principle.
The alternative, developed in detail later in this guide, is to choose a confidence level as a matter of policy and to compute the contingency that delivers it for each project individually. That turns contingency from a convention into a decision, and it makes the decision reviewable: a board can ask why this project is being funded at P80 while that one is at P60, and the answer will be about criticality and risk appetite rather than about habit. It also removes the perverse dynamic in which contingency is negotiated as a bargaining chip, because the number stops being an opinion.
Quantification sounds like a specialist activity requiring statisticians and dedicated software. In its essentials it is not. It requires four things: knowing which drivers matter, expressing each as a range rather than a point, representing which of them move together, and combining them by simulation rather than by arithmetic on single values. Everything else is refinement. The sections below take each step in turn, in the order a team would actually work through them.
A capital cost estimate may run to thousands of lines, but the outturn is usually governed by a few dozen. The first task is to find them, and the test is simple: a driver matters if a plausible movement in it produces a material movement in the total. A line item that is large but tightly bounded, such as a signed fixed-price package, contributes little variance. A line item that is modest but wildly uncertain, such as the extent of contaminated ground remediation, can dominate. Size and uncertainty both count, and it is their product that matters, not either alone.
Do this work with the people who own the estimate lines rather than in an analyst’s office. A risk workshop that asks each discipline lead what they are least sure about, and what would have to happen for their number to be badly wrong, surfaces drivers that no top-down review would find. It also does something quietly important for governance: it makes the uncertainty a shared property of the team rather than an external criticism of the estimate. Teams defend numbers they were asked to justify and improve numbers they were asked to bound.
Aim for a working list in the range of twenty to sixty drivers for a substantial project. Fewer than that and you are almost certainly hiding variance inside aggregate lines; many more and the model becomes hard to maintain and hard to explain, and the marginal driver contributes nothing to the answer. Resist the temptation to model every line of the estimate as uncertain. The goal is a model that a sponsor can understand and interrogate, not a replica of the cost breakdown structure with distributions attached to everything.
For each driver, the team needs a low, a most likely and a high, defined carefully. The most common error is to elicit ranges that are far too narrow, because people anchor on their point estimate and then step a comfortable distance either side of it. The correction is to ask the question differently: rather than “what is your range?”, ask “describe a plausible world in which this line comes in much higher, and tell me what that number is”, and then the same for lower. Forcing a causal story attached to each bound produces wider and better-calibrated ranges than asking for the bound directly.
Be explicit about what the bounds mean. A “high” that means “the worst I can conceive” is not the same as a “high” that means “I would be surprised one time in ten”, and mixing the two across a model corrupts the result. The convention we recommend is percentile-based: the low and high are values you would expect to be beaten roughly one time in ten in each direction, and the most likely is the mode. This makes the elicitation testable against outturn later, which is how a team becomes calibrated. The mechanics of turning these three numbers into a usable distribution are covered in our three-point estimation guide.
Ranges should be asymmetric wherever the underlying reality is asymmetric, and on capital projects it usually is. Quantities can overrun by more than they can underrun because there is a floor at the design quantity and no ceiling on rework. Durations can extend indefinitely but cannot compress below a physical minimum. Permitting can take years longer than planned and cannot take less than the statutory period. Forcing symmetric ranges onto asymmetric drivers is one of the quiet ways a model is made to produce a comfortable answer, and it is worth checking for explicitly during review.
The choice of distribution shape matters less than practitioners fear and less than novices hope. For most drivers, a triangular or a PERT-style distribution built from the low, most likely and high is entirely adequate, and the difference between them is small relative to the uncertainty in the ranges themselves. Spending a week arguing about lognormal versus beta while the ranges were elicited casually is optimising the wrong thing. Get the ranges right first; the shapes are a second-order refinement.
There are exceptions worth knowing. Drivers dominated by a low-probability, high-consequence event - a permit refusal, a contractor insolvency, a geological surprise - are better modelled as discrete events with a probability of occurrence and a conditional impact, rather than smeared into a continuous range. Drivers with genuinely unbounded upside, such as certain classes of claim exposure, are better represented by a distribution with a fat tail than by a triangle, because a triangle’s hard ceiling will understate the extreme. And drivers that are multiplicative through the model, such as a productivity factor, propagate differently from additive ones and should be modelled as factors rather than as absolute amounts.
Modern tools reduce this to a set of choices a project professional can make. Incertive is built so that a decision can be described in plain business language and the distributional machinery handled behind the scenes, which matters for capital projects specifically because the people with the best knowledge of the drivers are engineers, estimators and package managers, not modellers. If a quantification can only be produced by a specialist, it will be produced once, at sanction, and never refreshed - which defeats the purpose, since the value of the analysis compounds with re-running it.
Correlation is the step most often skipped and the one that most changes the answer. Its effect is easy to state: positively correlated drivers make the aggregate distribution wider, because they tend to move badly together rather than offsetting; independent drivers make it narrower, because their errors cancel. A model that assumes independence where correlation exists will produce a distribution that is too narrow and a contingency recommendation that is too small, and it will do so with a confident-looking S-curve that invites exactly the false comfort quantification was supposed to remove.
The practical approach does not require a correlation matrix estimated from data you do not have. Identify the common causes. If several drivers would all be affected by a hot regional labour market, model that market condition once as a shared factor and let it drive each of them, rather than assigning pairwise correlations. Common-cause modelling is easier to explain to a sponsor, easier to defend in review, and less likely to produce an inconsistent matrix. It also maps naturally onto how risk actually propagates: through shared conditions rather than through mysterious statistical affinity.
The same logic applies across the cost and schedule models. Prolongation costs are driven by the schedule outcome, so they must be sampled consistently with it rather than given their own independent range. Escalation depends on when the spend occurs, which depends on the schedule. Treating these as separate independent inputs is a common and material error, and it always biases the answer the same way: optimistic. If your model shows a tighter distribution than the base rates for your asset class would suggest, missing correlation is the first place to look.
With drivers, ranges and correlations in place, the computation is mechanical. A Monte Carlo simulation draws one value for each driver, respecting the correlation structure, computes the resulting total cost and completion date, records them, and repeats. After enough iterations - typically many thousands - the recorded outcomes form an empirical distribution that can be read directly. There is no closed-form mathematics to follow and no approximation to argue about; the answer is simply what the model does when you run it a great many times. Our guide to Monte Carlo simulation for business covers the method in general terms, and you can see it operate on a small model in the Monte Carlo calculator.
Two views of the output do most of the work. The histogram shows the shape: where the mass sits, how skewed it is, how far the tail runs. The cumulative curve, universally called the S-curve, shows probability directly: for any cost, the proportion of simulated outcomes at or below it. From the S-curve you read the answer to the board’s question by inspection. Find the sanctioned budget on the horizontal axis, read up to the curve, and the height is the probability that the project comes in at or under it. If that height is thirty percent, the organisation is being asked to approve a plan that fails seven times in ten.
It is worth being clear about what the simulation does and does not tell you. It does not predict your project; it characterises the range implied by your assumptions. If the ranges are too narrow, the output will be too narrow, and no amount of iteration fixes that - a point sometimes summarised as the model being a mirror rather than a crystal ball. This is precisely why the base rates matter: they are the external check on whether your model’s width is credible. A cost risk analysis that returns a P90 only eight percent above the base estimate, on a project type where the historical mean overrun is forty percent, is telling you something about your elicitation, not about your project.
Contingency is where quantification pays for itself, because it converts a perennial argument into a decision with a stated rationale. The argument usually runs like this: the project team wants more contingency because they will be held to the number; finance wants less because contingency is capital that could be deployed elsewhere and because unspent contingency is invariably spent; and neither party has any evidence, so the outcome is decided by seniority. A distribution replaces the argument with a question about risk appetite, which is a question the organisation is actually competent to answer.
The method is straightforward once the S-curve exists. Decide what probability of sufficiency the organisation wants for this project. Read the outturn cost at that probability from the curve. Contingency is the difference between that figure and the base estimate. Nothing about this is exotic, and it has the property that the resulting number is explainable in one sentence: this contingency funds the project to an eighty percent chance of coming in at or under budget, given the ranges the team provided.
Choosing the level is a governance decision rather than a technical one, and it should be made deliberately and consistently. A useful frame is to ask what happens if the project overruns. If an overrun would be absorbed comfortably from operating cash flow, a lower confidence level is rational and the capital freed can be deployed elsewhere. If an overrun would force a covenant breach, an equity raise, a cancellation of other investment, or a public inquiry, the appropriate level is much higher, because the cost of the tail is not proportional to its size. Guidance for government cost estimating has long emphasised that a credible estimate requires risk and uncertainty analysis and the identification of a range of confidence levels with adequate contingency and management reserve, rather than a single deterministic figure.
A required step, not an optional one - the GAO Cost Estimating and Assessment Guide treats risk and uncertainty analysis, the identification of confidence levels, and adequate contingency as characteristics of a reliable programme cost estimate.
One nuance often missed: the right confidence level is not the same at every level of aggregation. Individual projects in a large portfolio can reasonably be funded near the median, because overruns on some will be offset by underruns on others, and the portfolio as a whole can be held at a high confidence level with less total capital than funding every project individually at that level. This portfolio effect is real and material, but it only works if the projects are genuinely independent, which they often are not - a shared supply chain, a shared labour market or a shared economic cycle correlates them, and correlated portfolios need more capital, not less.
It helps to distinguish two pots with different owners and different purposes. Contingency covers the variability within the defined scope: the quantities, rates, durations and conditions that were always going to vary. It belongs to the project and is drawn down by the project manager against defined criteria. Management reserve covers changes to the scope itself: work the organisation decides it wants that was not in the sanctioned case. It belongs to the sponsor or the board and should require a decision to release.
Conflating them causes a predictable failure. If scope changes can be funded from contingency, then contingency is consumed early by discretionary additions and is unavailable when the variability it was sized for actually materialises. The project then appears to be running to budget for eighteen months and then overruns sharply, which is the characteristic profile of a project whose contingency was spent on scope. Keeping the pots separate makes the pattern visible early, because scope-driven draws appear against the reserve where the sponsor can see them.
The corollary is that contingency should have drawdown rules written before it is needed. A simple and effective rule is that contingency may only be drawn against a risk that was identified in the quantification, and that each draw must be recorded against the driver it relates to. This produces, almost as a by-product, the calibration data described later: at the end of the project you know which drivers consumed the contingency, and therefore which ranges were too narrow. Organisations that do this for three or four projects develop estimating judgement that no external benchmark can supply.
Schedule contingency deserves its own treatment because schedules behave differently from costs. Costs add: a saving on one package genuinely offsets an overrun on another. Durations along a path add, but parallel paths do not offset at all - the merge takes the latest. This means that where several parallel chains feed a milestone, the milestone is late if any of them is late, and the probability of all of them being on time is the product of their individual probabilities. Five parallel paths each with an eighty percent chance of hitting their date give the milestone roughly a one-in-three chance, not an eighty percent chance.
This effect, known as merge bias, is why deterministic critical-path schedules are systematically optimistic even when every individual duration estimate is honest. It also explains the familiar experience of a project that is on schedule at every gate and late at the end. A probabilistic schedule risk analysis captures the effect automatically, because the simulation takes the maximum of the merging paths on every iteration. A deterministic schedule cannot capture it at all, which is why adding a lump of float at the end of a critical-path schedule does not produce the protection people expect.
The practical consequences are worth stating. Buy schedule contingency where the merges are, not uniformly along the plan. Reduce the number of parallel chains feeding critical milestones where you can, because each one is an independent chance to be late. And treat schedule and cost contingency as linked rather than additive, since a late finish drives cost through prolongation, which means adding a full independent cost buffer on top of a schedule buffer double-counts. Our Monte Carlo simulation for project management guide goes through the merge mechanics in more depth.
A contingency that is set at sanction and never revisited is a static answer to a moving question. As the project proceeds, uncertainty resolves: the ground is investigated, the packages are let, the long-lead items arrive, the productivity rate becomes observable. Each resolution should shrink the remaining distribution, and therefore the contingency required to hold a given confidence level. A well-run project re-runs its quantification periodically and reports both the remaining contingency and the confidence level it currently buys, which is a far more informative status metric than percentage of contingency spent.
This produces the drawdown curve that experienced sponsors look for: contingency declining broadly in step with the resolution of risk. Two deviations are diagnostic. If contingency is being consumed faster than risk is resolving, the project is in trouble whether or not it is currently reporting to plan. If contingency is untouched deep into execution while significant uncertainty remains unresolved, it may be being hoarded, which has its own cost since the capital could have been released. Either pattern is visible only if the quantification is refreshed, which is the strongest practical argument for tooling a team can run itself.
It is also worth planning what happens when contingency is exhausted before the work is. The honest answer is that the project returns to the sponsor with a revised distribution and a request, and that this should be treated as a normal event with a defined process rather than as a failure requiring blame. Organisations that punish the first request for additional funding reliably train their projects to conceal the need for it until concealment is no longer possible, which converts a manageable overrun into a crisis. That dynamic is cultural, and it is covered in our guide to building a risk-aware culture.
A probability of success tells you where you stand. It does not tell you what to do. The output that drives action is the ranking of drivers by their contribution to the spread, conventionally displayed as a tornado diagram. This is the part of a quantitative risk analysis that most directly changes decisions, and it is the part most often left in an appendix.
A tornado diagram ranks drivers by how much moving each one across its range moves the outcome, holding the others at their central values, or by the statistical contribution of each to the variance of the output. The bars are sorted longest at the top, producing the funnel shape that gives the chart its name. The reading is immediate: the drivers at the top are the ones that determine whether the project lands well or badly, and the drivers at the bottom, however alarming they sound in a workshop, are not worth managing intensively.
The value of the chart is that it is frequently counter-intuitive. Teams devote attention to the risks that are vivid, recent, or personally threatening, and a ranked sensitivity often shows that the dominant driver is something mundane: a quantity growth allowance, a productivity factor, a currency exposure on a single large package. That mismatch between attention and influence is the ordinary state of affairs on capital projects, and correcting it is one of the highest-return things a quantification does. Our sensitivity analysis guide covers the methods in detail, and the tornado diagram feature page shows the output form.
One caution: a tornado reflects the ranges you supplied. A driver appears influential partly because it is genuinely volatile and partly because someone gave it a wide range, and an important driver given an artificially narrow range will not appear at all. Reviewing a tornado is therefore also a review of the elicitation. If a driver you know to be treacherous sits at the bottom of the chart, the correct response is to check its range rather than to relax about it.
The step from ranking to action is where analysis becomes management. For each of the top drivers, the question is what could narrow its range or shift its centre, and what that would cost. The available moves are more varied than teams assume. Investigation narrows epistemic uncertainty: more boreholes, a survey, a pilot, a prototype. Contracting shifts exposure: a fixed price, a target cost with a pain-gain share, an escalation cap, a hedge. Design reduces exposure: a simpler interface, a proven technology, a modular scope that can be staged. Scheduling reduces exposure: ordering the long-lead item earlier, removing a parallel path, decoupling a dependency.
Each of these has a price, and the discipline is to compare the price against the movement in the distribution it buys. A hedge that costs one million dollars and removes four million of downside variance is good value; the same hedge on a driver that ranks eleventh in the tornado is a waste. Without a ranked, quantified view, mitigation budgets get allocated by advocacy, and the loudest risk owner wins. With one, the allocation is defensible and the arguments are about the numbers rather than about the people.
It is equally important to identify the drivers you should not spend on. A driver with a wide range that no available action can narrow - a commodity price, a weather window, a regulatory decision outside your influence - should be carried rather than managed, and carried explicitly in contingency at the chosen confidence level. Attempting to manage genuinely aleatory variability consumes management attention and produces nothing. Distinguishing what to mitigate from what to fund is one of the more valuable habits a capital organisation can build.
The capability that most distinguishes quantitative practice from qualitative is that a proposed action can be tested in the model before it is bought. Change the range on the driver to reflect what the mitigation would achieve, re-run, and observe what happens to the probability of success and to the contingency required. If the probability moves from sixty-two to seventy-eight percent, the action is worth its price; if it moves to sixty-four, it is not. This turns mitigation planning from an act of faith into an evaluation, and it is quick enough to do live in a review meeting when the tooling supports it.
The same mechanism supports comparing structurally different options: a modular build against a single large facility, an early order against a later one, a lump-sum contract against a reimbursable one with a target. These comparisons are hard to make on a single-point basis because the options differ mostly in their variance rather than in their central case, and variance is invisible to a deterministic model. On a probabilistic basis the comparison is direct: which option has the higher probability of clearing the hurdle, and which has the worse tail. You can try a small version of this reasoning in the project success calculator.
Doing this well requires the model to be fast and re-runnable by the people in the room, which is a tooling requirement more than an analytical one. A model that takes a specialist two days to re-run will be re-run at sanction and never again, and the option-testing benefit will be lost entirely. This is one of the criteria we weigh most heavily in our comparison of project risk analysis tools, and it is the reason Incertive is built around describing a decision in plain language and getting a probability, a driver ranking and recommendations back in under a minute.
Capital governance is almost universally organised as a series of gates: a screening decision, a selection decision among options, a definition phase, a sanction or final investment decision, then execution reviews and finally handover and benefits realisation. The gate model exists because capital projects are irreversible, and gates are the mechanism by which an organisation preserves the option to stop while stopping is still cheap. Quantified risk should appear at every one of them, and it should say something different at each.
A gate is not a progress checkpoint. It is a decision point at which the organisation chooses whether to spend the next tranche of money, and the only question that matters is whether the case for continuing is stronger than the case for stopping or waiting. Gates degrade into rubber stamps when they are treated as milestones in a schedule rather than as decisions, and the symptom is that no project has ever been stopped at one. An organisation whose gates never reject anything does not have gates; it has a reporting cadence.
The reason gates decay is usually momentum. By the third gate a project has a team, a budget line, an executive sponsor and a place in the corporate narrative, and stopping it has become an admission rather than a decision. This is where a quantified probability changes the dynamic, because it makes the case for continuing testable. A project whose probability of clearing its hurdle has fallen from seventy percent to thirty-five between gates has produced evidence, and evidence is much harder to argue past than a feeling that things are getting harder. Our complete guide to go/no-go decisions develops this in general terms; the capital-specific point is that the gate is where the option value lives.
The analysis should mature with the project. Asking for a full quantitative cost risk analysis at screening is wasteful, and accepting a percentage contingency at sanction is negligent. A sensible progression looks like this:
The escalating rigour is deliberate. Early gates need speed and the outside view; late gates need detail and the inside view built properly. What must not happen is the common inversion, in which the early gates are decorated with spurious detail because a model was available, and the sanction gate is waved through on a percentage because the schedule was tight.
The range around the outturn is widest at the start and narrows as definition improves, which is the shape drawn in the first figure of this guide. Two things narrow it and one thing does not. Investigation narrows it, because it converts epistemic uncertainty into knowledge. Commitment narrows it, because a signed fixed price replaces a distribution with a number, at a price. What does not narrow it is time alone: a project that has been running for eighteen months without resolving anything has the same range it started with, minus the money it has spent.
This gives a practical test for whether a definition phase is working. At each review, ask which drivers have narrowed and by how much. If the answer is none, the front-end work is generating documents rather than knowledge. It also gives a way to price the front-end programme: each proposed study can be assessed by how much it would narrow the driver it addresses, and studies that would not move a top-ten driver can be deferred. Very few organisations manage their front-end this way, and those that do reach sanction with materially tighter distributions.
The cone also explains why the sanction gate deserves special weight. It is the point at which the largest remaining option is exercised, and after it the range narrows mostly through commitment rather than through knowledge. Everything that was not understood before sanction is either paid for in contingency or discovered in execution, and the discovery route is always more expensive. This is the analytical case for front-end loading restated in the language of distributions, and it is why the sanction paper is the single most important document in the capital process.
At sanction, three numbers should be on the page together: the base estimate, the contingency and the confidence level it buys, and the probability that the investment clears its hurdle rate given both cost and benefit uncertainty. Most approval papers carry only the first. Adding the second and third changes the character of the discussion, because it forces the committee to accept an explicit probability rather than an implicit one. Nobody who has been told a project has a thirty-five percent chance of coming in at budget can later claim to have been surprised by an overrun.
It also enables a more sophisticated set of decisions than approve or reject. A committee looking at a distribution can approve conditionally: release funding for long-lead items only, require a specific investigation before full release, cap exposure by staging the commitment, or approve at a higher contingency with a defined drawdown regime. These middle options are how experienced capital organisations preserve optionality, and they are only available if the risk has been quantified, because each of them is a statement about which part of the distribution the organisation is willing to accept.
Finally, the sanction record should preserve the model, not just the number. When the project completes, the ability to compare outturn against the distribution that was actually approved is what converts a single project into organisational learning. Storing only the approved figure loses the information that mattered. Our page on project risk tolerance covers how organisations set and record the confidence levels they are prepared to fund, which is the policy layer above the individual sanction decision.
The method is universal; the drivers are not. A sensible model starts from the drivers that actually dominate in the asset class being built, and the base rates that apply to it. The following notes are not exhaustive, but they capture where the variance usually sits in each of the main capital-forming sectors.
In building construction the dominant drivers are usually quantity growth against the design, subcontractor pricing and availability, productivity, ground and existing-structure conditions, and the pace of design completion relative to construction. The last of these is worth emphasising because it is so often the true root cause: construction that starts against an incomplete design generates rework and change orders whose cost is recorded against the construction packages, which disguises the fact that the driver was design maturity. Modelling design completeness explicitly, rather than burying it in a change allowance, produces a much more honest picture.
The base rates here are unforgiving, with only around a third of surveyed projects landing within ten percent of budget and large projects running dramatically over. Weather and site access add variance that is genuinely aleatory and belongs in contingency rather than in a mitigation plan. Escalation matters over long builds and interacts with the schedule. We treat the construction case at length in scenario planning software for construction projects, which works through the driver set and worked examples by project type.
One sector-specific caution: construction contracts allocate risk in ways that make the client’s exposure less than the project’s total but rarely as small as the contract implies. Claims, variations, and the possibility of contractor distress all leak exposure back to the client, and the leak correlates with the conditions that make the transfer valuable. Modelling a lump-sum package as a fixed number with zero variance is almost always wrong; modelling it as a tight distribution with a modest probability of a significant claim is closer to the truth.
Energy and utility capital programmes combine long lead times, heavy regulation, large single items of plant, and grid or network interfaces that are outside the project’s control. The EY analysis of the world’s largest power and water projects, cited earlier, found most of them late and most of them over, with hydro, water, coal and nuclear projects performing worst. The hydropower record is worse still, with the 245-dam study finding average real-terms overruns around ninety percent and average build times over eight years.
The drivers that dominate are usually connection and consenting timing, long-lead equipment slots, commissioning and grid compliance, and where the technology is first-of-a-kind, performance risk on ramp-up. Because the plant items are large and few, the outturn distribution is often lumpy rather than smooth: it is driven by a handful of discrete events rather than by the accumulation of many small variations. That argues for event-based modelling of the top items and continuous ranges for the balance, which is a hybrid most tools support and most spreadsheets do not.
Benefits risk in this sector is unusually tractable because output is metered and prices are observable, but it is also unusually exposed to policy. A project sanctioned on one tariff regime and commissioned under another has a benefits case that never existed at sanction. Where policy risk is material, it is better modelled as a scenario branch with an explicit probability than as a range on a revenue line, because the outcomes are discrete rather than continuous.
Transport is the best-documented sector and therefore the one where the outside view is easiest to build. The 258-project study gives asset-class base rates directly: roads around twenty percent, bridges and tunnels around thirty-four percent, rail around forty-five percent. Underground work carries the widest ground-conditions variance of any common construction type, and rail carries the additional burden of systems integration and possession access, which is why its base rate is the worst of the three.
Public infrastructure adds two drivers that private projects mostly lack: political timing and benefit optimism. Political timing changes the schedule in discrete jumps at elections and spending reviews. Benefit optimism, as the rail demand evidence shows, is frequently larger in magnitude than the cost overrun and receives a small fraction of the scrutiny. A public capital analysis that quantifies only cost is measuring the smaller half of the problem.
The other feature of public capital is that the appraisal framework often mandates an optimism-bias correction, which is a crude but effective form of reference class forecasting. Where such guidance applies, the right approach is to treat the mandated uplift as a floor and the project-specific quantification as the means to justify anything lower, rather than treating the uplift as a substitute for analysis. HM Treasury’s guidance is explicit that adjustments may be reduced as more reliable project-specific estimates and risk work are built up, which is exactly the incentive structure you want.
Process plant projects are dominated by scope definition quality, equipment lead times, module and fabrication yard performance, construction labour availability at what is often a remote site, and commissioning and ramp-up to nameplate. The ramp-up driver is distinctive: a plant that is mechanically complete on schedule but takes eight months rather than two to reach design throughput has destroyed a large slice of the benefits case without ever appearing as a cost overrun. Modelling the ramp explicitly, as a duration and a yield curve, catches value erosion that a pure capital-cost model misses entirely.
This is also the sector where front-end loading discipline is most developed and most clearly linked to outcomes, with the megaproject research pointing directly at poor front-end definition as a leading cause of failure. Practically, that means the definition-phase deliverables should be treated as gate criteria with teeth: a defined percentage of engineering complete, a site investigation closed out, an execution plan with named resources, and long-lead procurement strategy settled. Projects that pass this gate on a promise rather than on evidence are the ones that populate the failure statistics.
Remote and brownfield sites deserve separate treatment. Remote sites concentrate logistics, camp, and labour-availability risk, and these correlate strongly with each other, which widens the aggregate. Brownfield work concentrates interface risk with live operations, where a tie-in window missed by a week can cost more than the entire package it belongs to. Both are cases where the common-cause approach to correlation described earlier is essential, because independent sampling would understate the tail badly.
Programmes whose cost is dominated by software, systems integration or digital infrastructure behave differently from physical construction in one crucial respect: the tail is much fatter. The research on roughly 1,500 IT projects found an average overrun near twenty-seven percent - lower than many construction base rates - but with about one in six turning into a black swan at around two hundred percent cost overrun. A distribution with that shape cannot be summarised by an average, and contingency set from an average will be catastrophically wrong in the cases that matter.
The governance response is different too. Where the tail is fat, the priority is not a bigger buffer but a structure that limits exposure: smaller increments, earlier working software, hard stop-loss criteria, and a genuine willingness to terminate. This is a case where the correct application of quantified risk is to change the shape of the project rather than to fund its variance. Our analysis of why software projects fail covers the failure patterns and what distinguishes the programmes that avoid them.
Many modern capital projects are hybrids: a physical asset with a substantial control, data or digital layer whose integration determines whether the asset performs. These deserve to be modelled as two coupled programmes rather than as one, because their risk profiles differ so sharply. Treating the digital scope with construction-style contingency percentages consistently understates it, and treating the physical scope with technology-style pessimism wastes capital.
Analysis that nobody acts on is an expensive hobby. The governance layer determines whether a quantification changes what gets approved, how it gets funded and when it gets stopped. Most organisations have the analytical capability to do this and lack the reporting discipline to make it matter, which is why the same overruns recur in organisations that already own risk software.
Three roles are involved and they should not be held by the same person. The project team owns the base estimate and the ranges, because it holds the knowledge. The risk or assurance function owns the method: it challenges the ranges against base rates, checks that correlations are represented, and confirms that the model does what it claims. The sponsor or investment committee owns the confidence level and therefore the contingency, because that is a decision about the organisation’s risk appetite rather than about the project.
Where these collapse into one, the analysis loses its independence in a predictable direction. A project team that also sets its own confidence level will choose one that makes the project approvable. A risk function that also produces the estimate has no one to challenge it. An investment committee that also owns the ranges has effectively marked its own homework. The separation is not bureaucracy; it is the mechanism that keeps the numbers honest when the incentives are pulling the other way.
A fourth party is worth adding for large commitments: an independent review that is not part of the delivery chain at all, tasked specifically with challenging the outside view. The question it should ask is deliberately blunt: what did projects like this actually cost, and why should we believe this one will be different? Very few teams answer that question well, and the discomfort of being asked it is itself a useful corrective.
The business case is where cost risk and benefits risk finally meet, and it should be probabilistic on both sides. Rather than a single net present value or internal rate of return, the case should report the probability that the investment clears the organisation’s hurdle, together with the shape of the downside. Two projects with identical expected returns and different variances are not the same investment, and a deterministic case cannot tell them apart.
This reframing also improves capital allocation across a portfolio. When every case reports a probability of clearing the hurdle, projects become comparable in a way that expected values alone do not permit, and the committee can construct a portfolio with a chosen aggregate risk profile rather than accepting whatever profile emerges from approving the highest-return items one at a time. It also surfaces the projects whose attractive central case rests on a fragile assumption, which are precisely the ones that damage a portfolio.
A practical note on presentation: report the probability, the drivers and the recommended actions on one page, and put the model behind it for those who want it. Investment committees are not short of data; they are short of decision-shaped information. A case that says “sixty-one percent probability of clearing the hurdle, driven principally by scope growth and productivity, improved to seventy-four percent if the site investigation is completed before sanction” tells a committee what to do. Forty pages of model output does not. See a sample analysis for the shape of report we mean.
Board reporting on capital projects is often a traffic-light table with a commentary, which conveys sentiment rather than exposure. A probabilistic alternative fits in the same space and says far more. For each project report the current probability of delivering within the sanctioned cost and date, the confidence level the remaining contingency buys, the top three drivers, and the movement in each of these since the last report. Movement is the most informative element, because a project whose probability has fallen ten points in a quarter needs attention even if it is still nominally green.
The aggregate view matters as much as the project view. A portfolio in which every project is individually funded at the median will overrun in aggregate more often than a board expects, and a portfolio whose projects share a labour market or a supply chain is more correlated, and therefore riskier in total, than the individual figures suggest. Reporting portfolio-level exposure explicitly, including the correlation assumption behind it, is one of the more valuable things a capital function can provide to a board.
Language discipline helps here. “On track” should mean something specific and checkable, such as at or above the confidence level the project was funded at. “Amber” should mean the probability has dropped below a defined threshold. Without those definitions the status colours drift toward optimism as reporting dates approach, which is the well-documented watermelon effect: green on the outside, red within. Tying status to a computed probability removes the discretion that makes the drift possible.
The most valuable and least practised habit in capital risk management is recording what actually happened and comparing it against what was forecast. For each completed project, record the sanction distribution, the outturn, and which drivers consumed the contingency. Over a handful of projects this produces something no external benchmark can: a reference class specific to your organisation, your contractors, your regions and your estimating culture.
Calibration also measures the estimators rather than only the estimates. If a team’s stated P10 to P90 ranges contain the outturn far less than eighty percent of the time, the ranges are too narrow and the correction is systematic rather than project-specific. This is a measurable, improvable skill, and teams that receive feedback on it improve quickly. Teams that never see the comparison never improve, which is a large part of why the base rates have proved so stable over decades.
The organisational requirement is modest: a standard record kept at sanction and completed at handover, held somewhere durable. The obstacle is rarely technical. It is that the people who would have to record the comparison have moved on by the time the outturn is known, and nobody owns the loop. Assigning that ownership explicitly, usually to the assurance function, is a small governance change with a compounding return.
The following example is entirely hypothetical and every figure in it is illustrative. Its purpose is to show the shape of the analysis and the way the outputs change a decision, not to suggest that these are typical values for any real project.
A manufacturer proposes a new production line in an existing facility. The estimate presented to the investment committee is 48 million dollars (illustrative), with an 18-month schedule to mechanical completion and a further 3 months to full rate. Contingency is set at the corporate convention of 10 percent, giving a total ask of 52.8 million dollars (illustrative). The business case shows a net present value of 21 million dollars (illustrative) at the corporate hurdle rate, based on a ramp to full rate within three months and a five-year demand forecast.
Presented this way the decision looks easy. The return is comfortably above the hurdle, the contingency matches policy, and the schedule is consistent with previous lines. Every number in the paper is a single figure and none of them carries a stated confidence. The committee has, in effect, been asked to approve a distribution it cannot see.
A definition-phase quantification identifies eleven drivers that matter. The largest are the mechanical and electrical installation quantities, which are estimated from a design that is roughly 60 percent complete (illustrative); the installed cost of two long-lead machines whose quotations are eight months old; site labour productivity in a region with a tight market; the duration of the tie-in windows into live operations; and the ramp-up period to full rate. Each is elicited as a low, most likely and high with a causal story attached to each bound.
Two common causes are identified. Regional labour tightness affects installation productivity, the tie-in durations and the ramp-up support requirement simultaneously, so those three are driven by a shared factor rather than sampled independently. Global equipment lead times affect both machine packages together. Representing these two common causes rather than assuming independence widens the aggregate distribution materially (illustrative), which is exactly the effect described earlier.
The schedule model is linked to the cost model so that prolongation, extended site supervision and escalation are computed from the sampled duration rather than given independent ranges. The ramp-up period is modelled explicitly with a yield curve, so that a slow ramp reduces the benefits stream rather than merely appearing as a schedule note.
The simulation returns a distribution whose median outturn is 53 million dollars (illustrative), above the 52.8 million being requested including contingency. The probability of delivering at or under the total ask is approximately 46 percent (illustrative). Funding to an 80 percent confidence level would require a total of 61 million dollars (illustrative), implying contingency of roughly 27 percent of the base estimate rather than the 10 percent proposed. The tornado ranks installation quantities first by a clear margin, followed by productivity, then equipment escalation, then tie-in windows.
The benefits side is equally informative. Carrying the ramp-up uncertainty through to the business case, the probability that the investment clears the hurdle rate falls from a certainty implied by the deterministic case to about 68 percent (illustrative), with the ramp duration contributing more to that spread than the capital cost does. This is a common and counter-intuitive result on production assets: the value at risk sits in the commissioning period, not in the construction estimate.
Faced with this analysis, the committee has options that a single-point paper never offered. It can complete the design to 90 percent before sanction, which the model suggests would narrow the installation quantity range enough to lift the probability of meeting a 55 million dollar budget from 46 to roughly 63 percent (illustrative) for a front-end spend of about 0.6 million dollars (illustrative) - a trade that pays for itself many times over in expected terms. It can re-quote the long-lead machines with an escalation cap. It can release funding for the machines only, deferring the balance to a second gate after the design completes.
It can also decide to fund the project at a lower confidence level with eyes open, which is a legitimate choice for an organisation that can absorb the overrun. The point of the analysis is not that the project should be rejected or that the contingency must be 27 percent. It is that the committee now knows what it is accepting, can compare this commitment against the others in the portfolio on the same basis, and can choose deliberately among approve, defer, stage and investigate rather than choosing between approve and reject on a number with no width.
The final move is the one most often skipped: record the distribution alongside the approval so that the outturn can be compared against it in two years. That single act is what turns this project into evidence for the next one, and it costs nothing but the discipline to do it.
None of the above requires a large team, but it does require tooling that the people who own the decisions can operate themselves. The market ranges from spreadsheet add-ins through specialist schedule risk packages to integrated decision platforms, and the right answer depends less on feature lists than on who will actually run the model and how often.
Spreadsheets remain the default home of capital estimates and will continue to be, but they are a poor host for quantification. Simulation in a spreadsheet requires an add-in, and the resulting model is fragile in exactly the ways spreadsheets are famously fragile: formula errors propagate silently, ranges are hard to audit, correlation structures are hard to express and harder to review, and version control is by filename. The consequence is that the model becomes the property of the one person who built it, and re-running it after they move on is a project in itself.
The deeper limitation is that spreadsheets encourage deterministic thinking by design. A cell holds a number. Making a cell hold a distribution is an addition to the tool rather than its native mode, and everything downstream - the summary sheet, the report, the comparison against the previous case - is built around single values. Our analysis of Excel forecasting limitations goes through this in detail. For capital projects specifically, the fatal issue is refresh frequency: a model that is painful to update will not be updated, and an un-updated model at gate three is worse than none, because it carries stale authority.
This is not an argument for abandoning spreadsheets. The estimate can and usually should live in one. It is an argument for keeping the quantification somewhere that a package manager can open, adjust and re-run in an afternoon, so that the probability reported at each gate reflects what is currently known rather than what was known at sanction.
Judge tools against the failure modes described in this guide rather than against a feature matrix. The criteria that matter most for capital work are these:
It is also worth running one practical trial before committing: back-test the tool on a project whose outcome you already know. Feed in the ranges you would honestly have used at the time and see whether the analysis would have flagged the driver that actually caused the overrun. A tool that would have surfaced it earns your trust. One that would have blessed a project you now know was doomed has told you something useful about itself. Our comparison of project risk analysis tools and our overview of risk analysis software both work through the category in more depth, and the project risk analysis software page sets out how Incertive approaches it.
Introducing quantified capital project risk is a change to how money gets approved, not a software installation. The rollouts that survive follow roughly this sequence.
The sequencing matters more than the timing. The common failure is to buy a tool, train a specialist, and produce a beautiful analysis that arrives after the decision was made. Anchoring the pilot to a real gate date forces the analysis into the decision path, and being in the decision path is the only thing that makes it stick.
Three patterns kill the discipline once it is established. The first is the specialist bottleneck: one person owns the models, every project queues for them, and the analysis stops being refreshed. The remedy is tooling accessible to the people who own the drivers, plus a deliberate policy that the assurance function reviews models rather than building them.
The second is the ceremonial analysis: a quantification is produced for the sanction paper, states a probability nobody discusses, and is never mentioned again. The remedy is to require that the probability appears in the recommendation itself and that any change in it since the last gate is explained. If a number can be ignored without anyone noticing, it will be.
The third is range gaming. Once contingency is computed from ranges, ranges acquire political value, and teams learn that a wide range yields a larger buffer while a narrow one gets a project approved. The remedy is calibration: compare stated ranges against outturn, publish the comparison, and hold estimators to their hit rate. Once ranges are scored, they stop being negotiating positions and start being forecasts, which is the whole point. Our guide to building a risk-aware culture covers the organisational side of making that stick.
The evidence on capital projects is remarkably consistent and remarkably old. Overruns are the normal case, not the exception; they have not shrunk with better software or better project management methods; and they show up across sectors, countries and decades with a regularity that rules out bad luck as an explanation. What has not changed is the way most organisations govern the risk: a register of hazards, a matrix of colours, a percentage contingency, and a single number on an approval paper. The method has been failing for seventy years, and it is still the default.
The alternative is not complicated. Establish the base rate before you look at the estimate. Express every driver that matters as a range with a causal story behind each bound. Represent the conditions that make drivers move together. Simulate rather than add. Read the probability of sufficiency off the S-curve and choose the confidence level you are funding to as a deliberate decision. Rank the drivers, spend mitigation money on the top of the list, and test each action in the model before you buy it. Re-run at every gate, and compare outturn against forecast when the project completes so the next estimate is better than this one.
None of that removes uncertainty, and none of it guarantees a project lands on budget. What it does is stop the organisation from accepting exposure it has not measured. A committee told that a project has a forty-six percent chance of coming in at the requested figure may still approve it, and may be right to. What it can no longer do is be surprised, and the difference between a managed exposure and a surprise is most of what separates a capital programme that compounds value from one that consumes it.
If you want the broader foundations behind this - the frameworks, the vocabulary and the practical routines for running risk on any delivery programme - go back to the pillar guide, project risk management for non-technical leaders. From there, the natural next steps are reference class forecasting for building the outside view, sensitivity analysis for ranking the drivers, and Monte Carlo simulation for the engine underneath it all.
You do not need a modelling team to begin. Describe a commitment in plain language and see the probability, the ranked drivers and the recommendations for yourself: try the go/no-go calculator on a gate decision you are facing, explore how the analysis is built on the platform, or get started and run your next capital paper as a distribution rather than a number. The next sanction decision is coming either way. Make it with the odds in front of you.
Capital project risk is the exposure that a large, mostly irreversible investment in a physical or long-lived asset will cost more, take longer, or deliver less than the case on which it was approved. It is not a list of hazards; it is a measurable quantity - the range of outturn cost, schedule and benefit around the sanction estimate, and the probability attached to each part of that range. Quantifying it means expressing every uncertain driver as a range rather than a single figure, simulating them together, and reading the result as a probability of meeting the approved budget and date.
You quantify capital project risk by replacing single-point estimates with ranges on the drivers that actually move the outturn - quantities, rates, productivity, escalation, durations, permitting timing, commissioning - then running a Monte Carlo simulation that samples those ranges thousands of times, respecting the correlations between them. The output is a distribution of possible outturns and a cumulative S-curve, from which you read the probability of landing at or under the sanctioned budget, the cost at any confidence level such as P50 or P80, and a ranked list of the drivers responsible for the spread.
There is no universal percentage, and a round-number rule such as ten percent is a guess dressed as a policy. The defensible method is to choose a confidence level and fund to it: run the quantitative risk analysis, read the outturn at that confidence level from the S-curve, and set contingency as the difference between that figure and the base estimate. A project the organisation must not fail might be funded at P80 or higher; a portfolio of many small projects can be funded closer to P50 because overruns and underruns offset across the portfolio. The number is then a decision about risk appetite that can be explained, rather than a convention nobody can defend.
A risk register identifies and tracks discrete threats and who owns them; quantitative risk analysis measures what those threats, plus the ordinary variability in every estimate line, do to the outturn as a whole. A register scored by probability times impact produces ordinal colours that cannot be added, compared across projects, or turned into a contingency figure. Quantitative risk analysis produces a distribution with real units - money and time - so it answers the questions a board actually asks: what is the chance we exceed the budget, how much would we need to be safe, and which drivers should we spend mitigation money on first.
Three causes compound. First, projects are sanctioned on estimates made when scope definition is still immature, so the true range around the number is far wider than the number implies. Second, estimating is systematically optimistic: HM Treasury instructs UK appraisers to apply explicit optimism-bias uplifts precisely because forecasts skew low, and independent research has found the pattern persists across decades and sectors. Third, risk is usually recorded qualitatively rather than measured, so no one ever computes the probability that the approved budget is enough. Quantifying the range at sanction, and re-running it at every gate, addresses all three.
Incertive quantifies capital project risk without a modelling team. Describe the commitment in plain language and get a probability of success, the ranked drivers behind it, and the changes that most improve your odds - in under 60 seconds.
Analyze My ProjectBack to Blog