Module 5 — When is a prediction good?
Lesson 9 of 14
Probabilistic forecasts
5.9 Probabilistic Forecasts
Tomorrow's electricity demand will be:
42.3 GW
Will it?
Perhaps.
But what does that number actually mean?
Maybe demand will be 42.3 GW.
Maybe it will be 41.8 GW.
Maybe an unexpected cold snap pushes it above 45 GW.
Maybe a major industrial facility shuts down and demand falls.
The future has not happened yet.
So while 42.3 GW might be our best estimate, pretending that we know the future to one decimal place hides something extremely important:
uncertainty.
A probabilistic forecast tries to represent that uncertainty explicitly.
Instead of asking only:
What do we think will happen?
we ask:
What are the different things that could happen, and how likely do we think each of them is?
From point forecasts to probabilistic forecasts
A point forecast gives us a single predicted value.
For example:
[ \hat{y}=42.3\text{ GW} ]
This might be the expected electricity demand at 18:00 tomorrow.
A probabilistic forecast instead describes a distribution of possible outcomes.
For example, our model might tell us:
[ P(D < 40\text{ GW})=0.05 ]
[ P(40 \leq D < 42\text{ GW})=0.25 ]
[ P(42 \leq D < 44\text{ GW})=0.45 ]
[ P(44 \leq D < 46\text{ GW})=0.20 ]
[ P(D \geq 46\text{ GW})=0.05 ]
Now we know much more.
The model still thinks demand somewhere around 42–44 GW is most likely.
But it is also telling us that substantially higher and lower outcomes remain possible.
Instead of compressing everything the model knows into one number, we retain information about the range of possible futures.
The future is a distribution
This is an important conceptual shift.
We often talk about the future as though there is a future state waiting to be discovered.
But from the perspective of a decision being made today, there are many possible future states.
We might represent them as:
[ y_1,y_2,\ldots,y_n ]
with associated probabilities:
[ p_1,p_2,\ldots,p_n ]
such that:
[ \sum_{i=1}^{n}p_i=1 ]
Our forecast therefore becomes something like:
[ P(Y) ]
rather than simply:
[ \hat{Y} ]
The distinction is profound.
A point forecast says:
This is what will happen.
A probabilistic forecast says:
Given what we currently know, this is what we believe could happen, and this is how likely we think the different possibilities are.
That is a much more honest description of prediction.
Expected values
Once we have probabilities associated with possible outcomes, we can calculate an expected value.
Suppose tomorrow's demand could take three simplified values:
| Demand | Probability |
|---|---|
| 40 GW | 20% |
| 42 GW | 50% |
| 46 GW | 30% |
The expected demand is:
[ E[D]
(40)(0.2)+(42)(0.5)+(46)(0.3) ]
[ E[D]=42.8\text{ GW} ]
So our point forecast might simply report:
Expected demand: 42.8 GW
But notice what has happened.
We have compressed an entire distribution:
[ {40,42,46} ]
with probabilities:
[ {0.2,0.5,0.3} ]
into one number:
[ 42.8 ]
That number is useful.
But information has been lost.
The expected value does not tell us how uncertain the forecast is.
Two forecasts can have the same expected value
Consider two forecasts.
Forecast A
There is almost no uncertainty:
[ P(D=42\text{ GW})=1 ]
Forecast B
There is considerable uncertainty:
[ P(D=32\text{ GW})=0.5 ]
[ P(D=52\text{ GW})=0.5 ]
Both have the same expected value:
[ E[D]=42\text{ GW} ]
But they describe radically different situations.
In Forecast A, planning around 42 GW may be straightforward.
In Forecast B, planning around 42 GW could be disastrous.
Demand will actually be either much lower or much higher.
This illustrates why uncertainty itself can matter for decision-making.
Prediction intervals
One way to communicate forecast uncertainty is through a prediction interval.
Instead of saying:
Tomorrow's demand will be 42.3 GW.
we might say:
Expected demand is 42.3 GW, with a 90% prediction interval of 39.8–45.7 GW.
Conceptually:
[ P(39.8 \leq D \leq 45.7)=0.9 ]
The interval gives us a range within which we expect the outcome to fall with some specified probability.
A narrower interval represents greater certainty.
A wider interval represents greater uncertainty.
For example:
90% interval: 42.0–42.6 GW
suggests a very different forecasting situation from:
90% interval: 32–53 GW.
The central prediction might be identical.
The uncertainty is not.
Quantile forecasts
Another useful approach is to predict quantiles.
Suppose an electricity-demand model produces:
[ Q_{0.1}=39.5\text{ GW} ]
[ Q_{0.5}=42.1\text{ GW} ]
[ Q_{0.9}=45.8\text{ GW} ]
The median forecast is:
[ Q_{0.5}=42.1\text{ GW} ]
while the 10th and 90th percentiles describe lower and upper parts of the predicted distribution.
We can interpret these approximately as:
- 10% of outcomes are expected below 39.5 GW,
- 50% below 42.1 GW,
- 90% below 45.8 GW.
Quantiles can be particularly useful when different decision-makers care about different parts of the distribution.
Someone planning ordinary operation might care about the median.
Someone responsible for system security might care much more about the 95th or 99th percentile.
The tails matter
Suppose a flood model predicts:
[ P(\text{major flood})=0.01 ]
A 1% probability sounds small.
But if a major flood would cause billions of pounds of damage and threaten thousands of lives, it cannot simply be ignored.
Likewise:
[ P(\text{electricity shortage})=0.005 ]
might represent a rare event.
But electricity systems cannot necessarily be designed solely around the most likely outcome.
The consequences of the unlikely outcome may be enormous.
This is why the tails of probability distributions matter.
A forecast is not merely about identifying the centre of the distribution.
Sometimes the most important information is contained at its edges.
Probability and consequence are different
Suppose there are two possible events.
Event A
Probability:
[ P(A)=0.5 ]
Cost if it occurs:
[ £10 ]
Event B
Probability:
[ P(B)=0.01 ]
Cost if it occurs:
[ £10,000 ]
Event B is dramatically less likely.
But its consequences are dramatically larger.
A decision-maker therefore cannot look at probability alone.
A simplified expected cost calculation gives:
[ E[C_A]=0.5\times10=£5 ]
while:
[ E[C_B]=0.01\times10,000=£100 ]
The less likely event creates the larger expected cost.
This is one of the bridges from prediction towards decision-making.
A forecast tells us what might happen.
A decision system must consider:
[ \text{probability} + \text{consequence} + \text{objectives} + \text{constraints} ]
to decide what to do about it.
Weather forecasts are probabilistic for a reason
Weather provides perhaps the most familiar example.
Consider:
70% chance of rain tomorrow.
The forecast is not saying:
It will rain for 70% of the day.
Nor necessarily:
70% of the area will experience rain.
It is communicating uncertainty about whether a defined rainfall event will occur under the forecast conditions.
Tomorrow eventually produces only one observed outcome:
[ \text{rain} ]
or:
[ \text{no rain} ]
But before tomorrow happens, uncertainty exists.
That uncertainty is exactly what the probability attempts to represent.
And as we saw in the previous lesson on calibration, we can subsequently ask whether events assigned approximately 70% probability actually occur approximately 70% of the time.
Probabilistic forecasting and calibration therefore fit naturally together.
Electricity systems are forecasting machines
Electricity systems provide an especially useful example because enormous numbers of uncertain quantities must constantly be forecast.
Tomorrow's:
- electricity demand,
- wind generation,
- solar generation,
- temperature,
- network loading,
- generator availability,
- battery state,
- electricity prices,
- EV charging,
- industrial consumption
are not known perfectly in advance.
Consider wind generation.
A point forecast might say:
[ \hat{W}_{18:00}=12\text{ GW} ]
But perhaps the actual forecast distribution suggests:
[ P(W<8\text{ GW})=0.10 ]
[ P(8\leq W<12\text{ GW})=0.35 ]
[ P(12\leq W<16\text{ GW})=0.40 ]
[ P(W\geq16\text{ GW})=0.15 ]
A system operator now has much richer information.
The question is no longer simply:
How much wind do we expect?
It becomes:
How should we operate the system given the range of wind generation that might actually occur?
That is a decision problem under uncertainty.
Forecast uncertainty changes with time
Our uncertainty about an event usually changes as we approach it.
Imagine forecasting electricity demand.
One year ahead
There may be substantial uncertainty about:
- economic activity,
- weather,
- consumer behaviour,
- new technologies,
- industrial demand.
One week ahead
We know more.
One day ahead
Weather forecasts improve.
One hour ahead
We have recent measurements of actual demand.
Five minutes ahead
The possible range of outcomes may be much narrower.
We can think of the forecast distribution evolving as information arrives:
[ P(Y_{t+h}\mid X_t) ]
where (h) is the forecast horizon.
Generally, the further into the future we attempt to predict, the greater our uncertainty becomes.
This is another reason why saying simply:
The forecast is 42 GW
can hide important information.
We should also ask:
Forecast for when?
Forecasts should change when information changes
Suppose this morning we estimate:
[ P(\text{rain tomorrow})=0.30 ]
Later, a weather front develops unexpectedly.
New satellite measurements arrive.
Our updated forecast becomes:
[ P(\text{rain tomorrow})=0.75 ]
This does not mean the first forecast was necessarily wrong.
The information available to us changed.
A rational forecasting system should update its beliefs when new evidence arrives.
In probabilistic terms:
[ P(Y\mid X_{\text{old}}) ]
becomes:
[ P(Y\mid X_{\text{old}},X_{\text{new}}) ]
This connects directly to the Bayesian reasoning introduced earlier in the course.
Prediction is not a one-off declaration about the future.
It can be a continuously updated estimate conditioned on everything we currently know.
Scenario forecasts
Sometimes representing the entire probability distribution is difficult.
Instead, decision-makers may consider a set of scenarios.
For example:
Low-demand scenario
[ D=38\text{ GW} ]
Central scenario
[ D=42\text{ GW} ]
High-demand scenario
[ D=48\text{ GW} ]
These can be useful for asking:
What would we do if each of these futures occurred?
But there is an important distinction.
A scenario is not necessarily a probabilistic forecast.
Unless probabilities are assigned, we do not know whether the scenarios are believed to be:
- equally likely,
- extremely unlikely,
- deliberately pessimistic,
- deliberately optimistic,
- or simply illustrative.
Scenarios explore possible futures.
Probabilistic forecasts attempt to quantify our beliefs about their likelihood.
Forecasts can be conditional
Predictions also depend on assumptions.
Suppose an economic model predicts:
Electricity demand will reach 400 TWh by 2035.
Hidden inside that statement may be assumptions about:
- economic growth,
- EV adoption,
- heat-pump deployment,
- electricity prices,
- industrial policy,
- population,
- efficiency improvements.
The forecast is therefore not simply:
[ P(D) ]
It is closer to:
[ P(D\mid X) ]
where (X) represents assumptions and observed conditions.
Change (X), and the forecast may change.
This matters enormously when interpreting long-term forecasts.
A forecast is not a prophecy.
It is an inference conditional on information, assumptions and a model.
Forecasts can change behaviour
There is an even stranger problem.
Sometimes publishing a forecast changes the thing being forecast.
Suppose an electricity-price model predicts extremely high prices tomorrow evening.
Battery operators see the forecast and decide to discharge during that period.
Consumers reduce demand.
Generators change their schedules.
Those actions increase supply or reduce demand.
The predicted price spike may disappear.
Was the forecast wrong?
Not necessarily.
The forecast contributed to changing the system.
Similar effects occur elsewhere.
A navigation system predicts congestion and redirects drivers.
The redirection changes congestion.
A financial model predicts a stock price increase.
Traders buy the stock.
Their purchases affect its price.
A recommendation system predicts what people will watch.
Its recommendations change what people watch.
These are reflexive systems:
[ \text{prediction} \rightarrow \text{behaviour} \rightarrow \text{changed reality} \rightarrow \text{new data} ]
The forecast is no longer merely observing the world.
It has become part of the system changing it.
Probabilistic forecasts are not decisions
Suppose an AI system predicts:
[ P(\text{machine failure tomorrow})=0.08 ]
Should we shut the machine down?
The forecast cannot answer that question by itself.
We need to know:
- the cost of shutting it down,
- the cost of failure,
- whether maintenance is available,
- whether another machine can substitute for it,
- how safety-critical the machine is,
- what level of risk is acceptable.
Two organisations could receive exactly the same probabilistic forecast and rationally make different decisions.
That is because:
[ \text{prediction}\neq\text{decision} ]
The prediction describes uncertainty about the world.
The decision introduces:
[ \text{objectives} + \text{constraints} + \text{preferences} + \text{consequences} ]
This distinction will become central in the next module.
AI should sometimes say "I don't know"
There is a cultural tendency to judge intelligent systems by how confidently they produce answers.
But uncertainty is not necessarily evidence of poor intelligence.
Sometimes uncertainty is the correct answer.
Consider two systems.
System A
The component will fail in 17 days.
System B
The median predicted failure time is 17 days, but there is substantial uncertainty. There is a 10% probability of failure within 5 days and a 90% probability of failure within 35 days.
System A sounds more decisive.
System B tells us much more.
The goal of intelligent prediction should not be to eliminate uncertainty by pretending it does not exist.
It should be to reduce uncertainty where possible and represent the remaining uncertainty honestly.
The deeper lesson
We began this course with a simple observation:
The future does not exist yet.
We have data from the past.
We observe the present.
We construct models.
And from those models we infer possible futures.
So the relationship is not:
[ \text{past} \rightarrow \text{model} \rightarrow \text{future} ]
as though the model reveals something already written.
It is closer to:
[ \text{past data} + \text{present information} + \text{model} \rightarrow \text{distribution over possible futures} ]
The future eventually becomes one realised outcome.
But when we make a decision, that outcome is not yet known.
That is why probabilistic forecasting matters.
It forces us to distinguish between:
what we expect
and
what we know.
A useful intelligent system should therefore not merely answer:
What is most likely to happen?
It should increasingly help us understand:
What could happen?
How likely are the different possibilities?
How uncertain are we?
and eventually:
Given that uncertainty, what should we do?
That final question takes us beyond forecasting.
It takes us towards decision-making under uncertainty.