Changing a cell tells you what happens if an assumption is wrong. It does not tell you how likely that is, which assumption matters most, or what the odds of missing your target actually are. What-if analysis software answers all three.
Every finance team runs what-if analysis. Someone asks what happens if the deal closes a quarter late, or if churn runs two points above plan, or if the vendor quote comes in fifteen percent high, and an analyst opens the model, changes a cell, and reads the new bottom line. That loop is so familiar it barely registers as analysis at all. It is simply what a model is for.
The loop is also where a great deal of corporate decision-making quietly goes wrong. Not because the arithmetic is incorrect, but because changing one cell answers a question nobody actually asked. The real question is almost never "what happens if sales are ten percent lower." It is "how likely are we to miss, by how much, and which assumption should I go and pressure-test first." A spreadsheet that recalculates on demand cannot answer any of those three, and what-if analysis software exists to close that gap.
This guide covers what what-if analysis actually is, starting with the three tools Excel ships and what each is genuinely good for. It then examines where deterministic what-if analysis breaks down, why scenario tables systematically understate risk, and what changes when the same model is run probabilistically. It closes with a practical buyer’s framework for choosing what-if analysis software and a workflow for folding it into an existing planning cycle without rebuilding the finance stack.
The argument throughout is not that spreadsheets are bad. They are the most successful analytical tool ever built, and almost every model worth simulating starts life in one. The argument is narrower and, I think, harder to dismiss: a single recalculated number is an answer to a question about arithmetic, and the questions that matter in planning are questions about probability. Those require different machinery.
What-if analysis is the practice of changing the inputs to a model in order to observe the effect on its outputs. That definition is deliberately broad, because the practice is broad. It covers an analyst nudging a growth rate, a treasurer testing a covenant against a rate rise, an engineering leader asking what a two-week slip does to a launch date, and a board member asking what happens to the runway if the raise takes six months instead of three.
What unites these is a common structure. There is a model, which is some set of relationships between inputs and outputs. There are inputs whose values are uncertain. And there is at least one output that somebody cares about enough to defend in a meeting. What-if analysis is the act of exploring the space between the inputs and that output.
The interesting question is how much of that space you explore, and how you decide which parts are worth exploring. This is where practice diverges sharply, and where the difference between a spreadsheet and dedicated what-if analysis software becomes substantive rather than cosmetic.
Microsoft groups three distinct features under the What-If Analysis menu, and they are more different from one another than their shared home suggests. Understanding what each is for is the fastest way to see both the reach and the ceiling of spreadsheet-native what-if analysis.
Scenario Manager stores named sets of input values. Microsoft describes a scenario as a set of values that Excel saves and can substitute automatically in cells on a worksheet. You define Base, Upside and Downside, each specifying values for the same handful of cells, and switch between them. It is a bookmarking system for assumption sets.
Goal Seek runs the model backwards. It is the tool for when, in Microsoft’s phrasing, you know the result you want from a formula but you are not sure what input value the formula requires. You specify the output you need and the single input you are willing to move, and Excel solves for it. What break-even volume clears the hurdle rate; what price holds margin at forty percent.
Data Tables sweep a range. They apply when you have a formula that uses one or two variables, and they produce a grid of outputs across every combination of the values you supply. A one-variable table is a column of answers; a two-variable table is the familiar cross-tab with one input down the side and another across the top.
These three tools are well designed and genuinely useful, and any serious discussion of what-if analysis software has to start by acknowledging that a competent analyst with them can answer a great many questions quickly. The limitation is not craftsmanship. It is that all three are deterministic: each produces exact outputs for exact inputs, and none of them carries any notion of how likely any of those inputs were in the first place.
Consider what a two-variable data table actually tells you. Suppose you sweep unit volume from eighty thousand to a hundred and twenty thousand across the columns, and unit cost from eleven to fifteen dollars down the rows. You get a grid of, say, twenty-five margin figures. Every one of them is arithmetically correct.
Now ask the question the grid was built to inform: are we going to hit the margin target? The grid cannot say. It shows you that some combinations clear the target and others do not, but it attaches no weight to any cell. The corner where volume is lowest and cost is highest sits in the table with exactly the same visual prominence as the centre, despite being, in most real businesses, dramatically less likely. The table presents a uniform landscape over a decidedly non-uniform reality.
This is the central limitation, and it is worth stating precisely because it is so often blurred. A data table enumerates possibilities. It does not weight them. Any judgement about likelihood has to be supplied by the reader, from memory, in a meeting, under time pressure. That is a poor place to do probability.
The most common form of what-if analysis in corporate finance is the three-case model: Base, Upside, Downside. It is nearly universal, it is easy to explain to a board, and it is a genuine improvement on a single number. It also has a structural flaw that is rarely discussed.
The flaw is that the cases are usually constructed by moving every assumption in the same direction at once. The downside case has low volume and high cost and slow collections and a delayed launch. The upside has the mirror image. What this produces is not a range of plausible futures but two comparatively improbable ones bracketing a central estimate, because the probability of every variable landing on the same side of its distribution simultaneously is much lower than the probability of any one of them doing so.
The practical consequence is a peculiar kind of false comfort. The downside case looks alarming enough to feel like diligence has been done, while being unlikely enough that nobody really plans for it, and the genuinely probable outcomes, the ones where two assumptions disappoint mildly and a third surprises pleasantly, are never modelled at all. Those middling combinations are where most real outcomes live.
We cover the mechanics of building cases that are actually comparable in more depth in our guide to scenario analysis software, and the finance-specific version of the same problem in financial scenario planning software.
Deterministic what-if analysis fails in four specific, identifiable ways. None of them is a failure of the spreadsheet as a technology, and none is fixed by a more careful analyst. They are consequences of the method.
A data table handles one or two variables. Real models have considerably more than two uncertain inputs. A mid-market annual plan might carry uncertainty in volume, price, churn, wage inflation, hiring pace, vendor costs, collection timing, and foreign exchange, and that is a conservative list before anyone mentions a product launch or an acquisition.
With eight uncertain inputs and just three levels each, the space of combinations runs to several thousand cells. Enumerating that by hand is not merely tedious; it is unhelpful, because the output would be a table too large to read and still carrying no probability weights. The combinatorial wall is why deterministic what-if analysis in practice collapses back to three named cases, regardless of how many variables the model actually contains.
What-if analysis software addresses this by sampling the space rather than enumerating it. Instead of every combination at fixed levels, it draws thousands of combinations at random from the input distributions, which converges on the shape of the outcome without requiring anyone to read a thousand-cell grid.
This is the limitation described above, and it deserves its own heading because it is the one most often waved away. A downside case is a statement about what could happen. It contains no statement about how likely that is, and without a likelihood the case cannot be acted on rationally.
Consider two businesses, each with a downside case showing a twelve-million-dollar shortfall. In one, the downside is a one-in-twenty tail event. In the other, it is closer to a one-in-three outcome. These demand entirely different responses: the first is a risk to note and monitor, the second is a plan that should probably be changed before it is approved. A deterministic model presents them identically, because the model has no vocabulary for the difference.
Probability is what converts a list of things that could go wrong into a ranked agenda of things to do something about. Any what-if method that cannot express it is limited to the first half of that sentence.
One-variable-at-a-time analysis implicitly assumes that the variables are independent, and in business models they very often are not. If demand falls, sales volume drops, but so does the pricing power that would let you protect margin, and so, eventually, does the collection speed of customers under the same pressure. Wage inflation and hiring difficulty tend to arrive together. Supply disruption raises input costs and lengthens lead times at once.
When correlated inputs are varied one at a time, the analysis systematically understates the tails. Each variable is examined in a world where the others are behaving normally, which is precisely the world that does not obtain when things go wrong. Real bad quarters are bad because several things went wrong together, and that is exactly the case that one-at-a-time analysis is structurally unable to produce.
Proper what-if analysis software models correlation explicitly, so that when the simulation draws a low-demand sample it also draws the weaker pricing and slower collections that tend to accompany it. The effect on the output is usually substantial and almost always in the direction of more downside risk than the deterministic version suggested.
There is a final problem that sits underneath the other three, which is that the spreadsheet carrying the analysis may simply be wrong. This is uncomfortable territory, but the research literature is unusually consistent on it. Reviewing fifteen years of studies, Raymond Panko concluded that spreadsheet errors are both common and non-trivial, and that of the remedies proposed, only cell-by-cell code inspection has been demonstrated to work.
The relevance to what-if analysis is direct. Every scenario you run inherits every error in the model. A what-if exercise conducted on a flawed model does not reveal the flaw; it propagates it into three cases instead of one, and lends the whole exercise an air of rigour that the underlying arithmetic has not earned. The more elaborate the scenario work, the more convincing the wrong answer becomes.
This argues for two things. First, for keeping the model that drives decisions as small and inspectable as it can be. Second, for tooling where the model structure is explicit and versioned rather than buried in a lattice of cell references that only its author can navigate. We go further into the specific failure modes in the limits of Excel forecasting and in our side-by-side comparison with spreadsheet workflows.
The case against deterministic planning is not theoretical. It is one of the better-documented findings in management research, and it is remarkably stable across decades, geographies and project types.
The most cited body of work here belongs to Bent Flyvbjerg, whose studies of megaprojects produced what he named the Iron Law of Megaproject Management: over budget, over time, under benefits, over and over again. His analysis found that nine out of ten megaprojects run over budget, a rate too consistent across sectors and countries to be explained by local incompetence.
The pattern holds in technology as firmly as in infrastructure. A joint study by McKinsey and the University of Oxford examined large IT projects and found systematic overrun in both cost and duration alongside a shortfall in delivered value.
45% over budget, 7% over time, 56% less value - McKinsey and the University of Oxford studied 5,400 large IT projects and found that, on average, they ran 45 percent over budget and 7 percent over schedule while delivering 56 percent less value than predicted. Seventeen percent went so badly they threatened the existence of the company.
Source: McKinsey Digital
The direction of the error is the part that matters for what-if analysis. If estimates were merely imprecise, they would miss in both directions and average out across a portfolio. They do not. They miss predominantly to the downside, which means the error is not noise but bias, and bias is not fixed by working harder on the same method.
Flyvbjerg and Alexander Budzier extended the analysis to a large sample of IT projects and found something more troubling than a shifted average. Writing in the Harvard Business Review, they reported that alongside a substantial average overrun, a meaningful minority of projects failed catastrophically rather than moderately.
One in six is a black swan - Across roughly 1,500 IT projects, Flyvbjerg and Budzier found an average cost overrun of 27 percent, but one project in six became a "black swan" with a cost overrun near 200 percent and a schedule overrun around 70 percent.
Source: Harvard Business Review
This finding is precisely the one that deterministic what-if analysis cannot surface. A three-case model with a downside at, say, twenty-five percent over budget looks prudent against a twenty-seven percent average. It is silent on the one-in-six chance of something eight times worse, because the downside case was constructed by moving assumptions to plausible-sounding limits rather than by asking what the distribution actually looks like out in its tail.
A distribution, by contrast, has a tail by construction. When you simulate, the extreme combinations appear in proportion to their likelihood without anyone having to imagine them in advance, which is the only reliable way to get a tail into a plan. Nobody brainstorms their way to a black swan.
It might be argued that this evidence concerns large projects and has limited bearing on routine planning. The counter-argument is that the conditions producing it are now the normal operating environment for most businesses. PwC’s annual survey of global chief executives has repeatedly found macroeconomic volatility among the threats CEOs cite most often, which is a statement about the variance of the inputs every corporate model depends on.
When input variance is high, the gap between a point estimate and a distribution widens, and the cost of ignoring it rises. That is the practical case for what-if analysis software: not that the world has become unknowable, but that the range of plausible outcomes has widened enough that reporting the midpoint alone is now actively misleading.
The alternative to asking "what if this input were different" is asking "what is the distribution of outcomes given everything I know about all the inputs." That reframing is the whole of probabilistic what-if analysis, and the machinery to answer it has been well understood since the 1940s.
The first change is in how assumptions are stated. Instead of entering a growth rate of twelve percent, you state what you actually believe: growth is probably around twelve percent, realistically no lower than six, and plausibly as high as eighteen. That is a three-point estimate, and it carries strictly more information than the single number it replaces while requiring no more knowledge than the analyst already had.
This is a subtle but important point. Analysts frequently object that they do not know the distribution of their inputs. In practice they almost always do, informally: they know which assumptions are solid and which are guesses, and they know roughly how wrong each could be. A single-point model forces them to discard that knowledge at the moment of entry. A range-based model captures it. Our guide to three-point estimation covers the elicitation techniques, and the PERT calculator converts optimistic, likely and pessimistic values into an expected value and a standard deviation.
Different inputs deserve different shapes. A delivery date that can slip badly but cannot arrive early is skewed, not symmetric. A vendor quote that is either accepted or rejected is not continuous at all. Choosing an appropriate probability distribution for each input is a modelling decision, and a good tool makes it an explicit one rather than hiding it behind a default.
Once inputs are ranges, the model is run thousands of times. Each run draws one value from each input distribution, respecting any correlations you have specified, and computes the output. Ten thousand runs produce ten thousand outcomes, and the collection of those outcomes is the answer.
This is Monte Carlo simulation, and its appeal is that it requires no new mathematics from the user. The model is the same model. The formulas are the same formulas. The only change is that each input is sampled rather than fixed, and the output is collected rather than read. Everything difficult is handled by repetition, which is the one thing computers are unambiguously good at. The Monte Carlo simulation guide works through the method in full, and the Monte Carlo calculator runs a simulation directly in the browser.
The number of runs matters less than people expect. Convergence is usually good by a few thousand iterations for the middle of the distribution, though the extreme tails need more, since by definition few samples land there. What matters far more than iteration count is whether the input ranges and correlations were stated honestly.
A deterministic model returns a number. A probabilistic model returns a distribution, and reading one is a skill worth developing because almost every useful question can be answered from it directly.
Where is the mass concentrated, and is it where the plan assumed? How wide is the spread, which is a direct measure of how much you actually know? Is it symmetric, or does it have a long tail in one direction, which tells you whether the surprises are likely to be pleasant or unpleasant? And the question the whole exercise exists to answer: what proportion of outcomes fall on the wrong side of the line that matters?
That last figure is the one to take into a decision meeting. "There is a thirty-eight percent chance we exceed this budget" is a sentence a board can act on. "The downside case is twelve million over" is not, because it contains no information about whether that case is a remote possibility or a coin flip.
Once you accept that an assumption is a range rather than a number, a second question follows immediately: what shape does that range have? This is the part of probabilistic modelling that intimidates newcomers most, and it should not, because in planning work the honest answer is usually one of about six shapes and the choice between them is guided by the character of the quantity rather than by statistics.
It is worth being clear about how much the choice matters. For the middle of the output distribution, it matters rather little: almost any reasonable shape with the right centre and spread produces a similar median. For the tails, it matters a great deal, because the tails of the output are built from the tails of the inputs. If a decision turns on the probability of a severe outcome, the shape deserves thought. If it turns on the central estimate, do not agonise.
The triangular distribution is defined by three numbers: a minimum, a most likely value, and a maximum. Its popularity in planning is not an accident of convention. It maps exactly onto the way people actually hold uncertainty in their heads, which is as a best case, a worst case and an expectation, and it requires no parameter that a business user would struggle to supply.
Its shape is crude, with hard boundaries and straight sides that no natural process produces. That crudeness is usually acceptable and occasionally an advantage, because the hard boundaries make a strong claim explicit: you are asserting that values beyond the endpoints are impossible. That is often false, and being forced to notice it is useful. If you would not bet against an outcome outside your stated range, the range is too narrow.
Use it where you have a genuine floor and ceiling and an informed view of the middle. Vendor quotes, task durations with a hard deadline, headcount plans constrained by an approved budget.
The PERT distribution takes the same three inputs as the triangular but weights the most likely value more heavily, producing a smooth curve rather than a peak. In practice this means it concentrates more probability near the centre and less in the neighbourhood of the extremes, which for most estimates is a better reflection of belief.
The distinction matters more than it sounds. Under a triangular distribution, values close to the stated minimum are not far off as likely as values a little below the mode. Under PERT they are considerably less likely. If an estimator supplies a wide range to be safe, the triangular shape rewards that caution by shifting real probability mass out toward the edges, which can make an output look more volatile than the estimator intended. PERT is the more forgiving choice when ranges are elicited from people rather than derived from data. The PERT calculator shows the resulting expected value and spread for any three-point estimate.
The normal distribution is symmetric, unbounded in both directions, and characterised by a mean and a standard deviation. It is the right default for quantities that are themselves aggregates of many independent small effects, which is what the central limit theorem describes: total monthly transaction volume across thousands of customers, aggregate error in a large estimate, measurement noise.
It is the wrong default for most individual planning assumptions, for two reasons. It is symmetric, and business quantities are frequently skewed: a project can overrun by a factor of two but cannot underrun by one. And it is unbounded, so a simulation will occasionally draw a negative headcount or a negative price unless the model truncates it. Where a quantity has a natural floor of zero and a long right tail, the lognormal distribution is usually the better description.
The uniform distribution treats every value between two bounds as equally likely. It is rarely a good description of a real quantity, and it is occasionally exactly the right one, because it is the shape that encodes having no information beyond the bounds themselves.
There is a discipline in using it deliberately. If an analyst genuinely cannot say which part of a range is more likely, a uniform distribution says so honestly, whereas a triangular distribution with a mode guessed at the midpoint manufactures information that does not exist. It also has a useful rhetorical function in review: an input modelled as uniform is visibly flagged as one nobody understands well, which tends to attract exactly the attention it deserves.
Some of the most consequential uncertainties in a plan are not continuous at all. A regulatory approval is granted or refused. A key contract renews or lapses. An acquisition closes or collapses. Modelling these as a range is a category error, and doing so quietly smooths away the very discontinuity that makes them important.
These belong in the model as discrete events with an explicit probability and a defined consequence: a seventy percent chance of approval, and if it does not come, a specific cost and delay. The resulting output distribution is frequently bimodal, with two clusters rather than one hump, and that bimodality is real information. It says the plan has two quite different futures and the average of them describes neither, which is precisely the insight a single expected-value number destroys.
This is also where deterministic three-case modelling most obviously fails, because a base case has to either include the event or exclude it, and either choice misrepresents a genuinely uncertain binary. Our glossary defines the distribution types in more detail, and the probability distribution feature shows how each is applied to an input.
Correlation was identified earlier as the failure that most reliably causes what-if analysis to understate risk. It deserves a fuller treatment, because it is both the highest-leverage refinement available and the one most often skipped on the grounds that it seems technical.
When two inputs are positively correlated, they tend to be high together and low together. The effect on the output distribution is to widen it, sometimes dramatically. Independent errors partially cancel: one input running high is often offset by another running low, and the aggregate lands near the middle. Correlated errors compound instead, because the bad draws arrive in company.
This is the mathematical statement of something every experienced operator already knows. Bad quarters are not bad because one thing went wrong. They are bad because demand softened, which squeezed pricing, which stretched collections, which strained cash, all at once and for the same underlying reason. A model that samples those four independently will almost never generate that quarter, and will therefore report a probability of it that is far too low.
The practical consequence is that ignoring correlation produces a confident, professional-looking underestimate of tail risk. It is the most dangerous class of modelling error, because the output does not look wrong.
The usual objection is that nobody knows the correlation coefficient between two planning assumptions, which is true and largely beside the point. The choice in practice is not between a precise coefficient and nothing. It is between a rough coefficient and an implicit assertion of zero, and zero is a specific claim that is usually false.
A three-level convention is enough for most planning work: weak, around 0.3, for quantities that share some common driver; moderate, around 0.5, for those that clearly move together; strong, around 0.8, for those that are close to two views of the same underlying thing. Assigning these from judgement, and writing down the reasoning, is a large improvement over an unexamined zero.
Where history exists, use it. If you have several years of monthly data on two quantities, the observed correlation is a far better starting point than a guess, and it often surprises people in both directions. The test of whether the effort was worthwhile is simple: re-run with and without the correlations and see whether the answer to the decision question changes. If it does not, the refinement can be dropped with a clear conscience. If it does, you have found something important.
A handful of correlation patterns recur across almost every commercial model, and knowing them shortens the work considerably.
The fourth of these deserves emphasis because it explains one of the most common and least understood failures in project estimating. Adding up task-level estimates, each of which is individually reasonable, produces a total that is systematically optimistic, because the estimates share a common bias and a common set of dependencies. Modelling them as correlated recovers the realistic total.
The distribution is the raw material. In practice, teams work with a small number of summary figures drawn from it, and using them well is largely a matter of knowing which question each one answers.
A percentile states the probability that the outcome lands at or below a given value. P50 is the median: half the simulated outcomes are better, half worse. P80 is the value that eighty percent of outcomes fall below. These are definitional, requiring no source and no argument.
Their use, however, is where most of the value sits. P50 is a forecast. P80 is a commitment. If you plan to the P50 you should expect to miss half the time, which is fine for an internal expectation and disastrous for a number given to a board or a customer. If you commit to the P80 you should expect to miss one time in five, which for most commitments is an appropriate level of prudence.
The distance between P50 and P80 is therefore not an academic quantity. It is the contingency the plan requires, derived rather than negotiated. Most organisations set contingency by convention, a flat ten or fifteen percent applied to everything, which necessarily over-provisions the predictable parts of the plan and under-provisions the volatile ones. A percentile-derived contingency allocates it where the variance actually is.
Plot the cumulative probability against the outcome and you get an S-curve: for any value on the horizontal axis, the curve reports the probability of landing at or below it. It is the single most useful chart in probabilistic planning, because it turns every budget conversation into a readable trade-off rather than an argument about whose number is right.
Read it in either direction. Start with a budget and read up to find the probability of staying within it. Or start with a confidence level and read across to find what that confidence costs. The second reading is the one that changes conversations, because it reframes the discussion from "is this number right" to "how much certainty are we buying, and are we willing to pay for it."
It also makes the shape of the risk visible in a way a table cannot. A steep curve means a tightly constrained outcome where extra contingency buys little. A shallow one means genuine uncertainty, where each additional increment of confidence is expensive. Those two situations call for different decisions, and the curve distinguishes them at a glance.
Expected value weights each outcome by its probability and sums the result. It is the right tool for comparing options that will be repeated many times, because over many repetitions the average is what you actually get. The expected monetary value calculator computes it for a set of weighted outcomes.
The trap is applying it to decisions that happen once. A gamble with a positive expected value and a meaningful probability of ruin is a bad bet for an organisation that cannot survive the ruinous branch, no matter how attractive the average looks. Expected value implicitly assumes you get to run the experiment enough times for the average to materialise, and a company betting its balance sheet on a single large programme does not.
This is why the full distribution matters more than any summary of it. Expected value hides the tail by construction, since it averages over it. For one-shot decisions the relevant questions are about the tail: how bad is the bad branch, how likely is it, and can we survive it. Only the distribution answers those.
Knowing the odds is useful. Knowing what to do about them is more useful, and that requires a different question: of all the uncertain inputs, which ones are actually driving the uncertainty in the output?
In almost every model of any size, a small number of inputs dominate the variance of the result and the rest barely register. This is not a coincidence but a consequence of how models compose: an input matters to the output in proportion to both how uncertain it is and how strongly the output responds to it. An input can be wildly uncertain and still irrelevant if the output barely depends on it, and an input can be tightly known and still critical if the output is highly sensitive to it.
The practical implication is that effort spent reducing uncertainty should be concentrated, not spread. Refining an assumption that contributes two percent of the output variance is close to wasted work, however satisfying the refinement feels. Refining the one contributing forty percent changes the decision.
This is the single most actionable output of what-if analysis software, and the one most often missing from spreadsheet practice, because computing it by hand across a dozen inputs is laborious enough that nobody does it.
The tornado chart ranks inputs by their influence on the outcome, widest bar at the top, which produces the funnel shape that gives it its name. Each bar shows how far the output moves as that input ranges across its plausible values, so the chart answers "what should I go and find out" in a single glance.
The top two or three bars are the agenda. If sales volume dominates, the highest-value next action is not another modelling pass but a conversation with the commercial team about pipeline quality. If vendor cost dominates, it is a firmer quote. The chart converts an analytical result into an assignment, which is why it tends to be the artefact that survives the meeting.
It is equally useful in the negative. Inputs at the bottom of the chart can be left alone with a clear conscience, and being able to say "this assumption does not matter, here is why" is a genuine time-saver in review cycles where every number attracts equal scrutiny by default. Our sensitivity analysis explainer covers the technique in more depth, the sensitivity analysis overview shows how it fits a planning process, and the tornado diagram feature shows the live version.
There is an important technical distinction here that buyers should understand, because tools differ on it and the difference is material. The simple approach varies each input across its range while holding the others at their base values, then records the swing. This is easy to compute and easy to explain, and it is the method behind most spreadsheet tornado charts.
Its weakness is the one we met earlier: holding the others fixed assumes independence and ignores interactions. If two inputs only matter jointly, a one-at-a-time analysis will report both as unimportant, because moving either alone does little. The more rigorous approach apportions the variance of the output across the inputs and their interactions, capturing effects that only appear in combination.
For most planning work the simpler method is adequate and its transparency is a real advantage. For models with strong interactions, particularly where several inputs feed the same constraint, it can mislead. A tool that is explicit about which method it uses is preferable to one that shows a tornado chart without saying how the bars were computed.
The argument so far has been general. It is worth making it concrete, because the difference between deterministic and probabilistic what-if analysis is easiest to see on a single decision followed all the way through. Every figure in this example is (illustrative): it is a constructed case chosen to show the mechanics, not a real company’s data.
A software business is deciding whether to open a second sales region. The proposal requires hiring eight people, opening an office, and running at a loss for a period before the region contributes. The board has asked one question: will this be contribution-positive within eighteen months? The finance team has built a model and the base case says yes, at month fourteen (illustrative).
The deterministic what-if analysis has already been done in the conventional way. An upside case reaches contribution at month eleven; a downside case at month nineteen, just past the deadline. The pack shows all three. The recommendation is to proceed, on the basis that two of three cases clear the bar.
The first step is to identify which assumptions are genuinely uncertain and state each as a range rather than a point. In this model there are six that matter (all figures illustrative).
Note what this step surfaces before any simulation has run. The ranges are not all equally wide relative to their centres. Deal size varies by roughly a third either way, while price realisation varies by much less. That asymmetry is information the three-case model had thrown away, because in a three-case model every assumption moves to its limit together regardless of how uncertain it actually is.
Three correlations are worth encoding here. Deal size and deals per rep are mildly negatively correlated, because larger deals take longer to close (illustrative coefficient around -0.3). Attrition and time to productivity are positively correlated, since a team losing people ramps more slowly. And price realisation is mildly positively correlated with deal size, as both reflect the strength of the local market position.
None of these coefficients is known precisely. All three are defensible as directional judgements, and each is written down with its reasoning so a reviewer can disagree with it specifically rather than in general.
Running ten thousand iterations produces a distribution of months-to-contribution rather than three numbers. The median lands at month fifteen, a month later than the deterministic base case (illustrative). That shift is itself informative and is a common result: because several inputs are skewed to the downside, the median of the simulated output sits worse than the output computed from the modal inputs. The base case was never the middle; it was the result of assuming every assumption landed on its most likely value simultaneously, which is itself an unlikely event.
The decision-relevant figure is the probability of clearing eighteen months, and in this example it comes out at 63% (illustrative). That is a materially different statement from "two of our three cases clear the bar." It says the proposal is more likely than not to succeed on the stated test, and also that roughly one time in three it does not, which is a risk a board can now price rather than infer.
The distribution also has a long right tail. The worst five percent of runs do not land at month nineteen, as the deterministic downside case suggested, but between months twenty-two and twenty-six (illustrative), driven by the combination of attrition and slow ramp arriving together. That combination is exactly what the three-case model could not generate, because it moved one variable set at a time.
The tornado chart ranks the six inputs by contribution to the spread. In this example, time to productivity dominates, followed by deals per rep, with office cost and price realisation contributing very little (illustrative).
That ranking converts the analysis into an agenda. Negotiating the office lease harder, which is where a surprising amount of the team’s energy had been going, barely moves the answer. Shortening the ramp does. The highest-value actions are therefore the ones that attack ramp time directly: hiring two people with regional experience rather than eight generalists, running enablement before the office opens, or seconding an experienced rep from the home region for the first two quarters.
Re-running the model with a shortened ramp assumption shows what those interventions are worth in probability terms. If reducing most-likely ramp from five months to four lifts the probability of clearing eighteen months from 63% to 74% (illustrative), the secondment has a quantified benefit that can be weighed against its cost to the home region. That is a conversation the original three-case pack could not support at all.
The output of the exercise is not a chart. It is a recommendation with its reasoning attached: proceed, but with the ramp interventions funded, and with the review gate set at month nine rather than month twelve, because the simulation shows that ramp progress by month nine is the single best early predictor of which branch of the distribution the region is on.
That last point is worth dwelling on, because it is a class of insight only available from the probabilistic version. The simulation does not merely report the odds; it shows which early observable most separates the good runs from the bad ones, which is precisely what a review gate should be timed around. A deterministic model has no mechanism for producing that, because it has no runs to separate.
The method generalises, but the questions differ by function, and so does the shape of the answer that is useful. A few of the recurring applications are worth setting out, because the framing of the question determines whether the analysis gets used.
The classic FP&A application, and the one where the gap between current practice and available practice is widest. A rolling forecast that reports a single revenue figure per quarter invites exactly one conversation, which is whether that figure is right, and that conversation is unresolvable because nobody knows.
Reporting a range with an attached confidence changes the conversation to something answerable: is the spread acceptable, is it widening or narrowing quarter on quarter, and which drivers are responsible for the width. A forecast whose range is narrowing over successive cycles is a forecast the business is learning to control, and that is a far more useful management signal than whether a point estimate was hit.
The practical entry point is to keep the existing forecast process and add ranges to the handful of line items that actually drive variance, which is usually new business, churn and one or two cost lines. Our guide to financial scenario planning software covers the finance-specific workflow in more depth.
Pricing is a natural fit because the uncertainty is concentrated in one place, the demand response, and because the outcome is highly sensitive to it. A deterministic pricing model asks what revenue looks like at a ten percent increase assuming a given elasticity. The probabilistic version asks what the distribution of revenue looks like across the plausible range of elasticities, which is the question the decision actually turns on.
The output is usually more nuanced than expected. It is common to find that a price rise improves the median outcome while also widening the spread, so the choice is between a better expected result and a more predictable one. That trade-off is invisible in a deterministic model, which reports only the median, and it is frequently the thing the executive team most wants to discuss.
Hiring plans carry several correlated uncertainties at once: time to fill, ramp time, attrition and wage inflation, all driven partly by the same labour-market conditions. This makes them a case where independent modelling is especially misleading, and where a modest amount of correlation work changes the answer substantially.
The useful output is rarely the total headcount cost, which is reasonably predictable. It is the probability of having the required capacity in place by a given date, which is the constraint that actually binds on the plans depending on it. Framing the question that way also makes the trade-off between hiring ahead of demand and hiring behind it into something quantified rather than temperamental.
This is the domain where the evidence for probabilistic methods is strongest and the practice is, in places, most mature. The pattern of overrun documented by Flyvbjerg and by McKinsey is a direct consequence of single-point estimating at task level and summing the results, and the correction is to model task durations and costs as distributions with the correlations between them made explicit.
The output most valued by programme boards is the S-curve, because contingency is the decision and the curve prices it directly. Setting contingency at the P80 rather than at a conventional percentage tends to move money from the predictable parts of a programme to the volatile ones, which is where it was always needed.
Investment gates and stage-gate reviews are decision points with a binary output and, usually, a badly specified basis. The question "should we proceed" is answered far better by a probability of success against the stated criteria than by a narrative supported by a base case, because it forces the criteria to be stated numerically in the first place.
The secondary benefit is comparability across a portfolio. If every gate produces a probability of success and a ranked list of drivers, the portfolio can be managed as a portfolio rather than as a series of individually argued cases. Our go/no-go decision guide covers the criteria side, and the go/no-go calculator produces a verdict from weighted criteria.
Treasury applications are among the most valuable because the consequence of the tail is discontinuous. Breaching a covenant is not a bad outcome on a continuum; it is a different regime with its own costs. That discontinuity makes expected-value reasoning inappropriate and probability-of-breach the only sensible framing.
A model that reports "expected headroom of four million" is answering the wrong question. The right one is "the probability of breaching in any quarter of the next four is eleven percent, concentrated in Q3, and driven mainly by collection timing," which is a statement that tells the treasurer both how worried to be and what to do first.
The market for what-if analysis software is broad and the labels are unhelpfully overlapping. Spreadsheet add-ins, enterprise performance management suites, planning platforms, dedicated risk tools and general-purpose modelling environments all claim the capability. The following criteria separate them on the dimensions that affect whether the output is actually trusted and used.
This is the first question, and it is diagnostic. A tool that samples every input independently will produce a distribution that is too narrow and tails that are too thin, and it will do so silently, which is the dangerous part. The output looks like rigour and quietly understates risk.
Ask how correlation is specified, whether it can be set between arbitrary pairs of inputs, and what the tool does by default. A default of zero correlation is defensible if it is stated plainly; a tool that never raises the question at all is one whose confidence intervals should not be relied upon for decisions of any size.
Simulation uses random sampling, so two runs of the same model will differ slightly. That is expected. What is not acceptable is variation large enough to change the conclusion, or an inability to reproduce a specific prior result at all.
Reproducibility matters more than it first appears, for a reason that has nothing to do with statistics. The number will be questioned. Someone will ask why the figure in this month’s pack differs from last month’s, and "the simulation is random" is not an answer that survives a board meeting. Tools that support a fixed random seed can reproduce any run exactly, which means a difference between two runs can be attributed to an actual change in assumptions rather than to sampling noise.
This becomes essential the moment you want to compare scenarios. If you change one assumption and re-run, you want the difference in the output to reflect your change and nothing else. Running both versions on the same random draws isolates the effect of the change, which is the only way a scenario comparison means what it appears to mean.
A probability that cannot be explained will not be used. The output needs to carry its reasoning: which inputs drove the result, which assumptions the model is most sensitive to, and what would have to change for the answer to change. A tool that emits a single probability with no accompanying structure has moved the black box rather than opened it.
The test is simple. Take the output into a room with someone who was not involved in building the model and see whether you can defend it. If the answer to "why is it thirty-eight percent" is "the model says so", the tool has not done its job, whatever the quality of its arithmetic. Our methodology page sets out the approach we take to this, and the platform overview shows the resulting output.
Capability is worthless if the tool is not used. The traditional objection to probabilistic methods has always been effort: defining distributions for every input, specifying correlations, configuring a simulation and interpreting the output is real work, and in a planning cycle under deadline it is the first thing cut.
This is the dimension on which the category has changed most. Modern tools can take a plain-language description of a plan, identify the uncertainties implicit in it, propose sensible distributions and run the simulation without the user hand-specifying a parameter. That does not remove the analyst’s judgement; the proposals still need review, and reviewing a proposed range is considerably faster than deriving one from scratch.
The relevant question when evaluating is how long the first useful answer takes, not how powerful the tool is at full tilt. A capable tool used once a year during budgeting is worth less than a simpler one consulted before every significant decision.
Most real decisions are choices between alternatives rather than verdicts on a single plan. Should we phase the rollout or do it at once; hire ahead of demand or behind it; build or buy. A tool that evaluates one plan in isolation leaves the comparison to be done by eye across separate outputs, which is exactly where the sampling-noise problem above becomes acute.
Look for first-class support for variants: the ability to define several versions of a plan, simulate them on common random draws, and see them ranked with the trade-offs made explicit. The useful output is not merely which option wins on the average but which is more robust, because those are frequently different options and the difference is the decision. Our plan variants feature is built around exactly this comparison.
This is the criterion most often overlooked and, over time, the most valuable. A tool that produces probabilities but never records outcomes cannot tell you whether its probabilities mean anything. Calibration is the property that matters: of the things you said were eighty percent likely, roughly eighty percent should have happened.
Without outcome tracking, a forecasting process can be confidently wrong for years without anyone noticing, because each individual miss is explicable after the fact. With it, systematic optimism becomes visible and correctable, which turns the tool into something that improves rather than merely something that computes. Ask any vendor how their tool measures its own accuracy, and treat the absence of an answer as informative.
What-if analysis is claimed as a capability by several quite different classes of product, and the labels do little to distinguish them. It is worth separating them by what they are fundamentally built around, because that determines what they are good at and what they will always be awkward at.
These bolt simulation onto Excel directly. You keep the workbook, mark certain cells as distributions rather than values, and the add-in runs the model repeatedly and collects the outputs. The appeal is obvious: no migration, no retraining on the modelling language, and the full expressive power of a spreadsheet.
The limitations are inherited rather than introduced. The model remains a spreadsheet, with the fragility and inspection difficulty that implies, and it remains as hard to review as it was before. Governance is whatever the file-sharing arrangement provides, which in most organisations means very little. Add-ins suit a strong analyst working on a bounded, well-understood model who needs probabilistic output without changing anything else.
EPM and CPM platforms are built around the planning process rather than around any individual model: consolidation, workflow, driver-based planning, a governed dimensional data model and a reporting layer. Scenario capability is usually present and usually deterministic, offering named versions of a plan that can be compared side by side.
They are the right answer to a different question from the one this guide is about. If the problem is that six regions submit budgets in incompatible spreadsheets and consolidation takes three weeks, an EPM platform solves it and a simulation tool does not. If the problem is that the consolidated plan is a single number with no stated confidence, the platform may consolidate faster without addressing it at all. Many organisations need both, and it is worth being clear which problem is being bought for.
These are built around simulation as the primary object: distributions, correlation structures, convergence diagnostics, sensitivity decomposition and tail statistics as first-class features. They are the most capable option for genuinely quantitative work and are standard in insurance, energy, mining and large capital programmes.
Their cost is specialist assumption. They generally expect a user who is comfortable choosing a distribution family, reasoning about convergence and interpreting variance decomposition, and in a finance function without that background they tend to end up operated by one person and trusted by nobody else. Capability concentrated in a single individual is a fragile asset.
A smaller category built around the decision rather than the model. The organising object is a choice between options, and the output is a recommendation with supporting probability and drivers rather than a model artefact. These tend to prioritise time to first answer and explainability over modelling depth.
They suit organisations where the binding constraint is that rigorous analysis is not happening at all, rather than that it is happening imprecisely. The trade-off is that a tool optimised for a decision will not replace a full planning model, and pretending otherwise leads to disappointment on both sides. Incertive sits in this category deliberately, which is why the platform is organised around comparing plan variants rather than around building a general-purpose model.
Python with the scientific stack, R, or a notebook environment will do everything described in this guide and a great deal more, with complete control and no licence cost. For teams with the capability, this is a legitimate and powerful option, particularly where the model needs to integrate with data pipelines that already exist.
The costs are the familiar ones for bespoke software: somebody has to maintain it, the analysis lives in code that most of the finance team cannot read, and the reporting layer has to be built rather than bought. Organisations frequently underestimate the maintenance tail, and a model that only one person can run has a governance problem regardless of how good the mathematics is.
The practical selection test is to name the binding constraint before looking at any product. If it is consolidation and process, buy planning software. If it is that quantitative risk work needs to be deep and defensible to a regulator, buy a dedicated risk tool. If it is that decisions are being made on single-point estimates and nobody has time to change that, buy for time-to-answer and explainability. And if the constraint is that one excellent analyst needs simulation on a model that already works, an add-in is the cheapest good answer.
The failure mode is buying the most capable tool available and discovering that capability was never the constraint. Adoption is, almost always, and a tool that produces a defensible answer in an afternoon will change more decisions than one that produces a superb answer in a fortnight.
A model that informs material decisions is a control, and it should be governed like one. This is the least discussed aspect of what-if analysis software and frequently the one that determines whether it survives contact with an audit, a financing round or a change of CFO.
The most common governance failure is not a wrong assumption but an orphaned one. A range entered eighteen months ago by someone who has since left, never revisited, still driving a headline number. Nobody can say where it came from, so nobody feels able to change it, and it acquires authority purely through persistence.
The remedy is an assumption register: each uncertain input carries its range, the basis for that range, the person accountable for it, and the date it was last reviewed. This sounds bureaucratic and takes very little time once established, and it converts the review meeting from a debate about the output into a targeted examination of the two or three inputs that the sensitivity analysis says actually matter.
It also makes staleness visible. An input last reviewed before a material change in the business is a flag, and flags of that kind are far more useful than a general instruction to sanity-check the model.
The question that most often embarrasses a forecasting process is not "is this right" but "why is this different from last month." Without versioning, answering it requires reconstructing a previous state of the model from memory, and the honest answer frequently turns out to be that several things changed at once and nobody tracked which.
A tool that versions assumptions and can reproduce a prior run exactly turns that into a mechanical question. The difference between this month and last decomposes into the specific assumption changes that caused it, which both answers the question and, more usefully, shows whether the forecast moved because the world changed or because someone adjusted an input.
This is where the reproducibility criterion discussed earlier stops being a technical nicety. A process that cannot distinguish a real change from sampling noise cannot be governed, because every movement is arguable.
A healthy review separates two questions that are easy to conflate: is the model correct, and are the assumptions reasonable. The first is a technical question about structure and arithmetic and is best answered by inspection, once, thoroughly. The second is a business question, is answered by different people, and needs answering every cycle.
Tooling that keeps the two visibly separate makes both easier. When the model structure is explicit and versioned rather than distributed through a lattice of cell references, a reviewer can check the logic without wading through the assumptions, and can challenge an assumption without fearing they will break the logic. Panko’s finding that inspection is the only demonstrated remedy for spreadsheet error is an argument for exactly this separation, because a structure nobody can read is a structure nobody will inspect.
The governance step most often missing is the one that closes the loop. Forecasts are made, outcomes occur, and the two are rarely compared in any systematic way, which means the process cannot improve and cannot be shown to be trustworthy.
Recording each forecast with its date, its stated probability and its eventual outcome takes minutes per decision and produces, after a year, something no methodology argument can: evidence about whether this organisation’s eighty-percent forecasts come in around eighty percent of the time. Where they do not, the direction of the miss usually points at a specific and fixable habit, most often ranges that are too narrow on schedule and about right on cost.
The obstacle to adopting probabilistic what-if analysis is rarely intellectual. Most finance leaders accept the argument readily. The obstacle is operational: existing models, existing reporting packs, existing board expectations and a calendar with no obvious room in it.
The failure mode of enthusiastic adoption is trying to convert the entire planning apparatus at once. It does not work, because the effort lands all at the start, the benefit arrives much later, and the first hard conversation about an unfamiliar output happens when everything is already in flux.
A better approach is to pick a single consequential decision with a genuinely uncertain outcome, run it both ways, and present both. The deterministic case is what the organisation expects; the distribution is the addition. Showing them together is persuasive in a way that argument is not, because the gap between "the plan is twelve million" and "there is a thirty-eight percent chance we exceed twelve million" is self-evident once seen.
Choose something that matters enough to be argued about but not so large that being early in the learning curve is expensive. A significant hire, a pricing change, a phased rollout, a build-versus-buy decision. Decision-focused templates are a reasonable starting point.
Probabilistic what-if analysis does not require abandoning the spreadsheet that took three years to build. The relationships in it are usually sound; what is missing is any representation of uncertainty in the inputs and any machinery for propagating it.
In practice the useful move is to identify the handful of assumptions that actually drive the result, which the sensitivity analysis will tell you, and model those probabilistically while leaving the rest of the structure alone. A model with eight uncertain inputs and two hundred deterministic ones is a perfectly reasonable object and far easier to maintain than a fully stochastic rebuild.
This also limits the review burden. Eight distributions can be discussed properly in a meeting. Two hundred cannot, and a model whose assumptions cannot be reviewed is not more rigorous than a simpler one; it is merely harder to check.
The reporting change is the one requiring the most care, because it touches how the organisation talks about performance. Replacing a familiar single figure with a distribution, without preparation, invites the reaction that finance has stopped committing to numbers.
The framing that works is that the commitment is still there, and is now explicit about its confidence. The plan figure becomes the P80 rather than an unlabelled point, and the P50 sits beside it as the expectation. Nothing has become vaguer; the level of confidence has simply been stated instead of assumed, and it is worth pointing out that the old single number always had an implicit confidence level, one that nobody had ever measured.
Expect the first few cycles to generate questions about method. That is healthy, and answering them is how the numbers acquire authority. A tool that can show which assumptions drove a given figure makes those conversations considerably shorter.
Whatever else is adopted, start recording what actually happened against what was forecast, from the first cycle. This is cheap at the time and impossible to backfill later, and it is the input to every subsequent improvement in the process.
After a handful of decisions the record begins to answer questions that no amount of methodology debate can settle. Are our ranges too narrow? Are we systematically optimistic on schedule and realistic on cost? Which categories of decision do we forecast well, and which do we consistently misjudge? These are empirical questions, and they become answerable the moment someone starts keeping score.
This is also the most reliable route to organisational trust in the method. An argument that probabilistic forecasting is better is a position. A record showing that eighty-percent forecasts came in around eighty percent of the time is evidence, and evidence survives changes of management in a way that position does not.
There is a moment, usually the first or second time a team runs a proper probabilistic analysis, when the output says something nobody wanted to hear. The plan everyone has been working toward has a fifty-five percent chance of hitting its stated target. What happens next determines whether the method takes root or is quietly abandoned.
It is entirely predictable and occasionally correct. Someone will argue the ranges are too wide, the correlations too aggressive, the tail overstated. Sometimes they are right, and the discipline of checking is healthy; an analysis that cannot survive challenge should not survive.
The way to keep the challenge honest is to make it specific. "The model is too pessimistic" is not actionable. "The upper bound on ramp time should be six months rather than nine, because we have never seen worse than six" is a claim with evidence attached that can be tested by re-running. Requiring challenges to name an input and a reason keeps the conversation productive and, often, results in a better model rather than a discarded one.
What should not be acceptable is adjusting inputs until the output reaches a desired number. That is not modelling; it is working backwards from a conclusion, and it destroys the value of the exercise while leaving its appearance intact. A tool that versions assumptions makes this visible, which is one of the quieter arguments for governance.
The productive response to a bad probability is to change the plan, which is the entire point of doing the analysis before committing rather than after. A fifty-five percent chance of success is not a prediction to be accepted; it is a starting position to be improved, and the sensitivity analysis says where the leverage is.
There are generally four levers. Reduce the uncertainty, by buying information: a firmer quote, a pilot, a reference check, a paid discovery phase. Reduce the exposure, by restructuring: phasing the commitment, adding a decision gate, splitting a single large bet into two sequential smaller ones. Add buffer, by funding contingency at a confidence level rather than by convention. Or change the target, if the analysis shows it was never realistic and the honest move is to say so before rather than after.
Each of these can be evaluated by re-running the model, which is what makes the approach different from a discussion about risk appetite. The question "is a two-month pilot worth it" becomes "the pilot costs this much and moves the probability of success from 55% to 71% (illustrative), so is that worth the price," and that is a question a management team can settle in an afternoon.
The most underused of those levers is restructuring, because it changes the shape of the risk rather than merely the size of it. A decision made once, in full, at the beginning, carries the whole distribution. The same decision split into a commitment now and a larger commitment at a gate carries only part of it, because the second half can be informed by what the first half revealed.
This is where simulation earns its keep a second time, because it can tell you where the gate should be. The useful gate is placed at the point where the most decision-relevant uncertainty has resolved, which is not necessarily a convenient calendar boundary, and identifying it requires knowing which early observable most separates the good runs from the bad. That is a question only a distribution can answer.
A plan with a well-placed gate frequently has a worse expected value and a much better risk profile than the same plan committed in full, and for a one-shot decision that is usually the better trade. Comparing those two structures as plan variants rather than arguing about them in the abstract is the fastest way to settle it.
It is worth saying plainly that a probability below comfort is not automatically a reason to stop. Businesses take odds-against bets deliberately and correctly, when the payoff justifies it, when the downside is survivable, or when the alternative of doing nothing carries its own risk that nobody has modelled.
What the analysis provides is not a veto but an informed position: we are doing this with roughly a one-in-three chance of missing the target, we know the main driver is ramp time, we have a gate at month nine, and we can absorb the downside. That is a defensible decision. The same decision made on a base case that happened to land above the line is the same bet taken without knowing it was a bet.
Every argument in this guide rests on an assumption that deserves examining directly: that a probability produced by what-if analysis software is worth more than a guess. That is not automatic. A tool can produce beautifully presented numbers that bear no relation to reality, and the only way to find out is to check.
The instinct is to judge a probabilistic forecast by whether it was right, and that test does not work. If you say an outcome has a seventy percent chance and it does not happen, you were not wrong. Seventy percent events fail to happen three times in ten, and a forecaster whose seventy percent predictions always came true would be badly miscalibrated in the other direction, systematically understating their own confidence.
A single probabilistic forecast is close to untestable. Only the set is testable, and the property being tested is calibration: across everything you called seventy percent, did roughly seventy percent happen? That question requires a record, which is why outcome tracking was listed earlier as a selection criterion rather than a nice-to-have.
The mechanics are undemanding. For each forecast, record the date, the stated probability, the criterion that would count as success, and eventually what happened. The criterion matters most and is most often skipped: "the project succeeds" is not resolvable, while "contribution-positive by month eighteen" is, and the discipline of stating it precisely improves the forecast before any outcome arrives.
After a dozen or so resolved forecasts, patterns emerge. Group them into bands and compare the stated probability with the realised frequency. Most organisations discover they are overconfident, in the specific sense that their ninety percent forecasts come in nearer seventy, which is the same overconfidence that produces ranges that are too narrow. Discovering that is worth more than any single forecast, because it is correctable.
Formal scoring rules exist for this. The Brier score, which is the squared difference between the stated probability and the outcome treated as one or zero, averaged across forecasts, is the standard and has the useful property of rewarding both accuracy and appropriate confidence at once. It penalises a confident wrong call more heavily than a hedged one, and it penalises hedging everything at fifty percent too, which prevents the obvious gaming strategy.
The organisational effect of keeping score is larger than the analytical one. In most companies, a forecast that misses is explained rather than counted: the market moved, the deal slipped, a competitor did something unexpected. Each explanation is individually plausible, and collectively they ensure nothing is ever learned, because no pattern can form.
A calibration record cuts through this without requiring anyone to be blamed. It does not ask whether a particular forecast was reasonable; it asks whether this team’s eighty percents have historically behaved like eighty percents. That is a question about process rather than about people, and it turns out to be far easier to discuss honestly.
It also converts the adoption argument from an assertion into evidence. A finance team that can show its stated confidence levels have tracked reality for six quarters has something no methodology debate can produce, and something that survives a change of CFO. That, more than any feature, is what makes a forecasting process durable.
Introducing probabilistic what-if analysis to a team that has not used it reliably produces the same five objections. Each contains something real, and each has a practical answer.
The most common objection and the least substantial, because the alternative being defended is a single number, which is a far stronger claim about the distribution than any range. Entering twelve percent asserts that twelve percent is certain. Entering a range from six to eighteen asserts much less and is therefore easier to defend, not harder.
What lies behind the objection is usually discomfort at being asked to make uncertainty explicit, because an explicit range can be argued with in a way that an unexplained point estimate cannot. That discomfort is a reason to do it, not to avoid it. The useful reframe is that nobody is being asked to know the distribution, only to state what they already believe with its uncertainty included.
Partly true, and the answer is to give them one, labelled. A board wanting a commitment can have the P80, and a board wanting an expectation can have the P50, and both are single numbers. What has changed is that the confidence attached to the number is now stated rather than assumed.
In practice boards respond better to this than finance teams expect, because the alternative they have been living with is a number that missed often enough to erode trust, with no framework for discussing why. A figure presented as an eighty-percent commitment that misses one year in five is a process working as described; a point estimate that misses half the time with no stated confidence looks like incompetence.
This objection is legitimate and should be taken seriously, because plenty of tools do deserve it. A probability with no accompanying explanation is not an improvement on a point estimate; it is a point estimate with better marketing.
The answer is to insist on tooling that shows its work: which assumptions drove the result, how sensitive the answer is to each, and what would have to change for the conclusion to flip. Simulation itself is not opaque. Given the same assumptions and the same random seed it produces the same answer every time, and every part of the calculation can be traced. Opacity, where it exists, is a product decision rather than a property of the method.
The strongest objection, and the one that has historically killed adoption. If probabilistic analysis adds a week to a planning cycle that is already compressed, it will be cut, correctly, because a late plan has real costs.
Two things answer it. First, scope: modelling six uncertain inputs rather than sixty takes hours, not weeks, and delivers most of the benefit because most of the variance sits in a few places. Second, tooling: the setup burden that made this expensive is the part that has changed most, and a tool that proposes ranges from a plain-language description turns a modelling exercise into a review exercise. The relevant benchmark is time to the first useful answer, not time to a perfect model.
This one inverts itself on inspection. High unpredictability is not an argument against probabilistic analysis; it is the condition under which point estimates are least defensible and ranges most informative. A business whose outcomes are tightly constrained loses little by reporting a midpoint. A volatile one loses almost everything.
What the objection usually means is that the team does not believe a forecast can be accurate, which is true and is a different claim. The purpose of a probabilistic forecast is not to be accurate in the sense of hitting a number. It is to be calibrated, in the sense that things given an eighty percent chance happen about eighty percent of the time. That is achievable in volatile businesses, and it is the only kind of accuracy that a forecast in an uncertain environment can honestly aim at. Our probabilistic forecasting guide covers the distinction in full.
What-if analysis software is being adopted during a broader reappraisal of what finance technology is for, and the pattern in the survey data is worth noting because it bears directly on how to evaluate these tools.
Artificial intelligence has arrived in the finance function at scale, but the reported results are notably more modest than the adoption figures.
59% adopting, 7% reporting high impact - Gartner’s 2025 AI in Finance Survey of 183 CFOs and senior finance leaders found 59 percent reported AI use in the finance function, roughly flat on the prior year. Only 7 percent reported a high or very high impact from it.
Source: Gartner
That gap between adoption and impact is the most interesting number in finance technology at the moment. It suggests the tools are being deployed against the wrong class of problem, and a second Gartner finding points at what that class is: 45 percent of CFOs say their AI investments lean toward productivity, while 20 percent say they lean toward decision quality.
Productivity gains in finance are real but bounded. Producing the same forecast in half the time is worth something; it does not make the forecast more likely to be right. If the underlying planning method reports a single number where the honest answer is a distribution, accelerating it simply produces the same misleading output sooner. Decision quality is the constraint, and it is the smaller share of investment.
This is not an argument against applying AI here, but it does suggest where to apply it. The historical bottleneck in probabilistic analysis has been setup: translating a plan into a structured model, identifying which quantities are uncertain, and proposing defensible ranges for each. That work is judgement-heavy, repetitive, and precisely the sort of thing language models do well.
The division of labour that works is therefore narrower than the marketing usually suggests. Use the model to read a plan and propose the structure: here are the uncertainties implicit in what you wrote, here are plausible ranges, here is how they likely relate. Then use simulation, which is deterministic given its inputs and fully auditable, to compute the answer.
This matters because the two halves have different failure modes. A language model asked directly for a probability will produce a confident and unverifiable number. A simulation given explicit assumptions produces a number that can be traced back to them and re-run when they change. The description is the part worth automating; the calculation is the part worth keeping rigorous.
The direction of investment supports the case that this is where attention is turning. Deloitte’s quarterly survey of chief financial officers has found both a near-universal view that AI matters and cloud-based planning technology at the top of the cost agenda, which is a reasonable proxy for where finance leaders expect the next increment of value to come from.
The opportunity in that spend is to buy decision quality rather than throughput. A planning platform that produces the same deterministic three-case output faster has automated an approach the evidence has been questioning for forty years. One that reports the odds, ranks the drivers, and keeps score of its own accuracy is doing something the old process could not do at any speed.
Probabilistic what-if analysis has its own characteristic errors, and they are worth naming because they tend to be made confidently.
The most common error by a wide margin. Asked for a plausible range, most people supply one considerably too tight, a finding robust across decades of research on overconfidence. The consequence is a simulation that inherits the overconfidence and launders it into a professional-looking output with unjustifiably narrow confidence intervals.
Two practical remedies help. First, ask for the extremes before the centre: what is the worst this could plausibly be, what is the best, and only then what is most likely. Anchoring on the midpoint first reliably compresses the range. Second, check the range against history. If the last four projects of this type overran by fifteen to forty percent, a range topping out at ten percent is not a forecast but an aspiration. Reference class forecasting is the systematic version of that check.
A simulation reporting a 37.4 percent probability invites more confidence than it deserves. That figure rests on input ranges that were themselves estimates, and reporting it to a decimal place implies a precision the inputs cannot support.
Round, and speak in bands. "Roughly a one-in-three chance of exceeding budget" carries the decision-relevant content without the false precision. The value of the analysis lies in distinguishing a one-in-three from a one-in-twenty, not in resolving thirty-seven from thirty-nine.
Running ten thousand iterations of a model with a structural error produces ten thousand wrong answers and a great deal of misplaced confidence. Simulation amplifies whatever it is given, and it is worth repeating Panko’s conclusion that the only remedy demonstrated to work on spreadsheet errors is systematic inspection.
Before simulating, check the model deterministically. Do the base-case numbers reconcile to something known? Does the output move in the expected direction when an input is nudged? Are the units consistent throughout? These checks take an hour and prevent the most expensive category of error.
The opposite failure is over-elaboration: assigning a distribution to every input on the grounds that everything is uncertain. Technically true, practically counterproductive. A model with sixty stochastic inputs is unreviewable, and an unreviewable model is not more rigorous than a simple one.
Let the sensitivity analysis discipline the scope. Model the inputs that drive the variance and fix the rest at sensible point values. The resulting model is smaller, faster, easier to explain and very nearly as accurate, because the inputs you dropped were the ones contributing almost nothing to the spread.
The final error is analytical rather than statistical. A histogram is not a recommendation. Executives do not want to interpret a distribution; they want to know what to do, and an analyst who presents a chart without a conclusion has transferred the hard part back to the audience.
Lead with the decision, support it with the probability and the drivers, and keep the distribution available for the people who want it. "I recommend we approve at the higher figure, because at the original number we have a one-in-three chance of overrunning, and the single biggest driver is vendor cost, which we can firm up in two weeks" is what the analysis was for. The chart is evidence, not the argument.
What-if analysis software is not a new category so much as a correction to an old habit. The habit is answering questions about an uncertain future with a single number, and it persists mainly because the tooling made it the path of least resistance for thirty years.
The evidence that the habit is costly is not seriously in dispute. Large projects overrun in one direction with a consistency that rules out chance. Tails are fatter than named downside cases suggest. Volatility in the inputs is the environment rather than the exception. None of that is fixed by more careful single-point estimating, because the problem is not carelessness but the method.
What changes with probabilistic what-if analysis is the question being asked. Not "what is the number," which has no honest answer, but "what is the range, how likely is each part of it, which assumption is driving the spread, and what would we have to believe for this to go badly." Those have answers, and the answers are actionable in a way that a recalculated cell is not.
If you are evaluating tools, the criteria that separate them are the ones set out above: correlation handling, reproducibility, explainability, time to a first answer, genuine variant comparison, and whether the tool keeps score of its own accuracy. That last one predicts long-term value better than any feature list, because it is the only one that forces the method to prove itself.
None of this requires abandoning the tools a finance team already relies on. The spreadsheet remains where most models are born, the planning platform remains where the process lives, and the change being argued for here is narrower than a replacement: state the assumptions you are least sure about as ranges rather than points, let the machine explore the combinations rather than naming three of them by hand, and report the odds alongside the number. That is a modest change in method and a large change in what the output is capable of supporting.
Incertive was built around that view of the problem. You describe a plan in plain language, it identifies the uncertainties, runs the simulation, and returns a probability of success with the drivers ranked and the assumptions visible. You can read more about how the method works, compare it with a spreadsheet workflow, see what it looks like for executives, or check the pricing.
What-if analysis software lets you change the inputs to a model and see the effect on the outputs. Spreadsheet versions such as Excel’s Scenario Manager, Goal Seek and Data Tables do this deterministically, producing one exact output per set of exact inputs. Dedicated what-if analysis software adds probability: instead of entering a single value per assumption, you enter a range, the tool samples thousands of combinations, and the result is a distribution of outcomes with the odds attached rather than a single recalculated number.
Excel groups three features under What-If Analysis. Scenario Manager stores named sets of input values you can switch between. Goal Seek works backwards, solving for the input value needed to reach a result you specify. Data Tables sweep one or two variables across a range of values and show the output for each combination. All three are deterministic: they enumerate possibilities but attach no likelihood to any of them, which is the main reason teams move to probabilistic tools as models grow.
They answer different questions. What-if analysis asks what the outcome becomes if inputs change, so the output is a value or a distribution of values. Sensitivity analysis asks which inputs matter most, so the output is a ranking, usually shown as a tornado chart. In practice they are complementary: the what-if run tells you the odds of missing your target, and the sensitivity analysis tells you which assumption to go and pressure-test first to improve them.
Usually not. The relationships in an established model are generally sound; what is missing is any representation of uncertainty in the inputs. The practical approach is to identify the handful of assumptions that actually drive the result, which a sensitivity analysis will show you, model those as ranges, and leave the rest of the structure as fixed values. A model with eight uncertain inputs and two hundred deterministic ones is a reasonable and reviewable object.
For the middle of the distribution, convergence is usually good by a few thousand iterations, and running more changes the headline numbers very little. The tails need more runs, because by definition few samples land there, so if a decision turns on a rare but severe outcome it is worth increasing the count. In practice the iteration count matters far less than whether the input ranges and the correlations between them were stated honestly, which is where most of the error in a simulation comes from.
Incertive turns a plain-language plan into a probability of success, a ranked list of the assumptions driving it, and the changes that most improve your odds - the same way every time, with the reasoning visible.
Try Incertive TodayBack to Blog