Module 3 — Uncertainty, probability and inference
Lesson 5 of 13
Expected values
Probability distributions describe many possible outcomes.
But when making decisions, we often want a simpler summary.
One of the most important summaries is the expected value.
The expected value is the probability-weighted average of all possible outcomes.
It answers a question like:
If this uncertain situation were repeated many times, what average outcome would we expect?
For a random variable X, the expected value is often written:
E[X]
or:
𝔼[X]
At its simplest:
EXPECTED VALUE = sum of each outcome × its probability
A simple example
Suppose a game has two possible outcomes.
- 50% chance of winning €100
- 50% chance of winning €0
The expected value is:
0.5 × €100 + 0.5 × €0 = €50
So:
E[X] = €50
This does not mean you will receive €50.
You will actually receive either:
€100
or:
€0
The expected value is a summary of the distribution.
Expected value is not necessarily a possible outcome
Consider a fair six-sided die.
Possible values are:
1, 2, 3, 4, 5, 6
Each has probability:
1/6
The expected value is:
3.5
But a standard die can never show:
3.5
This is an important point:
Expected value describes the centre of a probability distribution. It does not necessarily describe an outcome that can actually occur.
Why expected value is useful
Suppose a business expects the following demand tomorrow:
| Demand | Probability |
|---|---|
| 80 units | 0.2 |
| 100 units | 0.5 |
| 120 units | 0.3 |
The expected demand is:
80 × 0.2 + 100 × 0.5 + 120 × 0.3
which gives:
102 units
That number provides a useful summary of uncertain demand.
It may help with:
- staffing,
- procurement,
- inventory,
- capacity planning.
But whether the business should prepare for exactly 102 units is another question.
Expected value combines probability and consequence
This makes expected value useful for decision-making.
Suppose there are two possible outcomes:
- a small gain that is very likely,
- a large loss that is unlikely.
Expected value allows both:
likelihood
and:
magnitude
to enter the same calculation.
For example:
- 99% chance of gaining €10
- 1% chance of losing €1,000
The expected value is:
0.99 × €10 + 0.01 × -€1,000
The rare loss matters because its magnitude is large.
But expected value can hide risk
Now compare two options.
Option A
Receive:
€50 with certainty
Option B
Receive:
- €100 with probability 50%
- €0 with probability 50%
Both have expected value:
€50
But they are clearly not the same experience.
Option A has no uncertainty.
Option B has substantial uncertainty.
So:
Expected value does not describe risk by itself.
Two distributions can have the same expected value and very different spreads.
The average is not the distribution
This is one of the most important lessons in probability.
Suppose two electricity-demand forecasts both have:
E[D] = 40 GW
But:
Forecast A
Demand is usually between:
39 and 41 GW
Forecast B
Demand is usually between:
25 and 55 GW
The expected value is identical.
The operational problem is not.
If insufficient capacity is costly, Forecast B may require much more reserve.
Expected value is a compression
A probability distribution may contain:
- hundreds,
- thousands,
- infinitely many possible values.
Expected value compresses all of that into one number.
That is useful.
But compression loses information.
Specifically, expected value does not tell us:
- spread,
- skew,
- tail risk,
- multimodality,
- uncertainty.
This is why expected value should rarely be interpreted alone.
Weighted averages
Expected value is just a weighted average.
Suppose:
X = 10 with probability 0.2
X = 20 with probability 0.3
X = 30 with probability 0.5
Then:
E[X] = 10×0.2 + 20×0.3 + 30×0.5
The more probable an outcome is, the more influence it has on the expected value.
Expected value for continuous variables
For continuous random variables, the idea is the same.
Instead of summing a small number of discrete outcomes, we effectively average across a continuous probability distribution.
You do not need the calculus yet.
The conceptual point is enough:
The expected value is the centre of mass of the probability distribution.
Where the probability is concentrated more heavily, it pulls the expected value towards it.
Expected values can be shifted by extreme outcomes
Suppose most outcomes are around:
€50
but there is a very small probability of:
€10 million.
That extreme value can strongly affect the expected value.
This happens in:
- finance,
- insurance,
- entrepreneurship,
- disaster risk,
- rare-event systems.
Expected values can therefore be sensitive to tails.
Mean and expected value
When we calculate the average of observed data, we often call it the sample mean.
When we describe the theoretical average of a random variable, we call it the expected value.
Conceptually:
EXPECTED VALUE → average implied by the probability distribution
SAMPLE MEAN → average observed in collected data
With enough representative observations, the sample mean may approach the expected value.
Expected value and repeated trials
Suppose a fair coin pays:
€1 for Heads
and:
€0 for Tails
The expected value per flip is:
€0.50
You will not receive €0.50 on any individual flip.
But after many flips, the average payout per flip should approach €0.50.
This is one connection between probability and long-run behaviour.
The law of large numbers
The law of large numbers describes the tendency for averages from repeated observations to move towards the expected value as the number of observations grows.
For example:
after:
2 coin flips
you might observe 100% Heads.
After:
10 flips
perhaps 70%.
After:
100,000 flips
the proportion of Heads is likely to be much closer to 50%.
Likewise, the average payout approaches its expected value.
This is why expected value becomes particularly powerful in large repeated systems.
Insurance relies on expected values
Consider an insurance company covering thousands of customers.
For one individual, an accident is highly uncertain.
But across a large population, aggregate losses may become more predictable.
The insurer can estimate:
expected loss per policy
and combine that with:
- operating costs,
- capital requirements,
- risk margins.
The individual outcome remains uncertain.
The aggregate system becomes more predictable.
Casinos rely on expected values too
A casino does not need to know who will win the next game.
Individual outcomes are uncertain.
Instead, it designs games where the expected value favours the casino.
Across large numbers of repeated games, small expected advantages accumulate.
This illustrates a powerful principle:
You do not need to predict every individual event correctly if the expected behaviour of many events is stable.
Aggregation can reduce uncertainty
Suppose individual household electricity demand is difficult to predict.
One person might:
- cook,
- charge a car,
- turn on heating,
- leave home unexpectedly.
But across one million households, much individual variability may cancel out.
Aggregate demand can be much more predictable.
Expected values become particularly useful at scale.
But correlated events do not cancel as easily
Suppose every household responds to the same cold weather.
Now their demand changes together.
The individual uncertainties are not independent.
Aggregate demand may change significantly.
This is why dependence between random variables matters.
Expected values alone do not tell us how risks combine.
Expected demand and capacity
Suppose expected hospital demand tomorrow is:
100 beds
and the hospital has:
100 beds.
Is that enough?
Not necessarily.
If demand is:
- sometimes 90,
- sometimes 110,
then having exactly the expected number of beds means shortages occur frequently.
Capacity planning often requires considering:
upper parts of the distribution
rather than only:
expected demand.
Expected supply can be misleading too
Suppose a renewable generator has expected output of:
100 MW.
That does not mean it can reliably provide 100 MW whenever needed.
Perhaps:
- sometimes output is 0 MW,
- sometimes output is 200 MW.
Expected energy production and firm capacity are different concepts.
This distinction matters in electricity systems and many other shared-resource problems.
Expected value and reliability
Suppose average demand is below average supply.
That sounds reassuring.
But if:
high demand
sometimes coincides with:
low supply
the system can still fail.
So system reliability depends on the joint distribution, not simply:
expected supply - expected demand.
This is why planning from averages alone can be dangerous.
Expected profit
Businesses often use expected value to compare uncertain decisions.
Suppose Project A has:
- 80% chance of earning €1 million,
- 20% chance of losing €500,000.
Expected profit is:
0.8 × €1m + 0.2 × -€0.5m
This produces an expected monetary value.
But whether the company should take the project still depends on:
- available capital,
- downside tolerance,
- alternative opportunities,
- strategic objectives.
Expected value informs the decision.
It does not determine it.
Expected value versus utility
Suppose someone has:
€1,000
and faces a gamble:
- 50% chance of gaining €10,000,
- 50% chance of losing the entire €1,000.
The expected monetary outcome may appear attractive.
But losing all available money may be much more harmful than the equivalent gain is beneficial.
People often value gains and losses nonlinearly.
This motivates the concept of utility.
Instead of asking only:
What is the expected amount of money?
we may ask:
What is the expected value of the outcomes to the decision-maker?
Money is not always linear in value
Consider two people.
Person A has:
€1,000.
Person B has:
€10 million.
A loss of:
€1,000
has very different consequences for them.
So identical monetary outcomes may have different utility.
This becomes important in:
- economics,
- insurance,
- public policy,
- service design.
Expected utility
Decision theory often extends expected value into expected utility.
Conceptually:
EXPECTED UTILITY = probability-weighted value of each outcome to the decision-maker
This allows us to represent attitudes towards:
- risk,
- loss,
- benefit.
The key point is:
The numerical outcome and the value we place on that outcome are not necessarily the same thing.
Expected social value is harder still
Suppose a policy creates:
- €100 benefit for 1,000 wealthy people,
- €100 cost for 1,000 low-income people.
The net monetary expected value may be:
zero.
But is society indifferent between the two distributions?
Not necessarily.
We may care about:
- who gains,
- who loses,
- fairness,
- vulnerability.
Expected aggregate value can hide distributional effects.
Expected values can hide who experiences the outcome
Suppose an AI system reduces average waiting time from:
30 minutes
to:
25 minutes.
That sounds good.
But perhaps:
- most users improve slightly,
- one vulnerable group waits much longer.
Average expected performance has improved.
Service fairness may have worsened.
This is another reason averages should not be the only objective.
Expected cost
Instead of expected benefit, we can calculate expected cost.
Suppose:
- 99% chance of no failure with cost €0,
- 1% chance of failure costing €1 million.
Expected cost is:
€10,000.
This can help compare mitigation options.
If preventing the failure costs:
€2,000
then prevention may appear economically attractive.
But again, expected cost does not tell the whole risk story.
Catastrophic losses complicate expected value
Suppose an event has:
0.001% probability
but causes:
catastrophic loss.
Even if expected monetary loss appears modest, society may still choose to avoid the risk.
Why?
Because some outcomes are unacceptable even at low probability.
This introduces:
- risk constraints,
- safety constraints,
- precautionary principles.
Expected value is not the only way to design decisions.
Constraints can override expected value
Suppose an autonomous vehicle has two possible routes.
Route A minimises expected travel time.
But it has a small probability of violating a safety constraint.
Route B is slightly slower but much safer.
A well-designed system may choose Route B.
The objective is not simply:
minimise expected time
It is:
minimise expected time subject to safety constraints.
This distinction will become central in optimisation.
Expected value under different actions
Suppose we can choose between:
Action A
and:
Action B.
Each action creates its own distribution of possible outcomes.
We can calculate:
E[outcome | Action A]
and:
E[outcome | Action B]
This allows us to compare decisions.
Conceptually:
CURRENT STATE
↓
ACTION A → DISTRIBUTION A → EXPECTED VALUE A
ACTION B → DISTRIBUTION B → EXPECTED VALUE B
Decision theory often begins here.
But the action changes the distribution
This is important.
The probability distribution is not always something we passively observe.
Our actions can alter it.
For example:
Do nothing
may create one distribution of future system states.
Install backup capacity
creates another.
Give treatment
creates another.
Brake
creates another.
Decision-making is therefore partly about choosing between distributions over futures.
Prediction becomes action evaluation
A predictive model asks:
What is likely to happen?
A decision model asks:
What is likely to happen under each possible action?
Then we can compare expected outcomes.
So:
PREDICTION
becomes:
COUNTERFACTUAL PREDICTION UNDER ACTIONS
which becomes:
DECISION.
This connects expected value to causal reasoning.
Expected value in reinforcement learning
Reinforcement learning uses a related idea.
An agent considers:
- states,
- actions,
- rewards.
It tries to choose actions that maximise expected future reward.
Conceptually:
STATE
↓
ACTION
↓
POSSIBLE FUTURE REWARDS
↓
EXPECTED RETURN
The system chooses actions partly based on expected long-term consequences.
We will explore this later.
Immediate versus future value
Suppose a robot can take:
Action A
Gain:
10 units now
Action B
Gain:
0 now
but create opportunities for:
100 units later.
A purely immediate expected value might choose A.
A long-term decision system may choose B.
This introduces time into expected value.
Expected future value
For sequential decisions, we may care about:
expected cumulative value across many future periods.
For example:
reward now + expected reward later + expected reward after that...
The present decision can affect all of them.
This is why state and sequential decision-making matter.
Discounting
Future outcomes are often weighted differently from immediate outcomes.
A simple framework may use a discount factor.
Future rewards receive progressively smaller weights.
Reasons can include:
- time preference,
- uncertainty,
- opportunity cost.
Discounting becomes important in:
- economics,
- finance,
- reinforcement learning.
But the choice of discount rate is itself a value judgement.
Optimising expected value shapes behaviour
Suppose an AI system is told:
Maximise expected clicks.
It will search for actions that increase expected clicks.
If clicks are only a proxy for what we really care about, the AI may become extremely effective at maximising the wrong quantity.
So the question is not only:
Can we compute expected value?
It is:
Expected value of what?
The objective matters more than the mathematics
Suppose a hospital AI optimises:
expected patients treated per hour.
It may favour:
- simple cases,
- fast treatments.
That could improve the metric while disadvantaging complex patients.
The expected-value calculation may be mathematically perfect.
The objective may still be poorly designed.
This connects directly to service design and fairness.
The expected value reflects probabilities and values
Expected value requires two ingredients:
probability
and:
outcome value.
So if either is wrong, the expected value can be wrong.
For example:
poor probability estimates
lead to:
poor expected value estimates.
Likewise:
poorly chosen outcome values
lead to:
poor decisions.
Both sides matter.
Model uncertainty affects expected value
Suppose a model estimates:
10% probability of failure.
But we are highly uncertain about that estimate.
Perhaps the true probability could be:
2%
or:
30%.
The calculated expected cost may appear precise.
But the probability input is uncertain.
We should therefore remember:
An expected value is only as trustworthy as the distribution used to calculate it.
Expected value can change with new information
Suppose expected demand is:
40 GW.
Then a weather forecast changes.
New expected demand becomes:
45 GW.
The random variable did not change definition.
Our information changed.
So the distribution changed.
And therefore:
E[D]
changed.
Expected values are stateful summaries of current beliefs.
Conditional expected values
Often we calculate expectations given some information.
For example:
E[D | cold weather]
means:
Expected electricity demand given cold weather.
Or:
E[waiting time | Monday morning]
Or:
E[default | income, debt, repayment history]
Machine-learning predictions can often be interpreted as estimates of conditional expectations.
Regression and expected value
Suppose a model predicts house price from:
- size,
- location,
- age.
A regression model may effectively estimate something like:
Expected sale price given these features.
Conceptually:
E[Price | Features]
The actual sale price can still differ.
The model estimates the conditional centre of the distribution.
A prediction is often an expectation
When a model reports:
predicted demand = 40 GW
it may actually mean:
Given the available information, 40 GW is our estimate of the expected value.
That is different from saying:
Demand will definitely be 40 GW.
This distinction should be explicit when predictions are communicated.
Expected value and point forecasts
This is why point forecasts are so common.
The expected value provides one sensible single number.
But the point forecast hides the underlying uncertainty.
A better interface might report:
Expected demand: 40 GW
Likely range: 37–44 GW
Now we have:
centre + uncertainty.
Sometimes the median is a better point forecast
The expected value is not always the best single summary.
Suppose a distribution is highly skewed.
The mean may be pulled toward extreme values.
The median may better represent a typical outcome.
Which point estimate is best can depend on how prediction error is scored.
This will become important when we study loss functions.
Different loss functions produce different optimal summaries
Suppose a model is penalised using:
squared error.
The best prediction tends to be the:
expected value.
If it is penalised using:
absolute error.
The best prediction tends to be the:
median.
This is a deep connection between:
how we define error
and:
what the model learns to predict.
We will return to this later.
Expected value is therefore not neutral
When we choose to summarise uncertainty using an expected value, we are implicitly choosing a particular representation.
That representation may be appropriate.
But it prioritises one property of the distribution:
its probability-weighted centre.
Other properties disappear.
This connects back to our general principle:
Every representation preserves some information and discards other information.
Expected value and resource allocation
Suppose expected demand for a shared resource is:
100 units.
Should we provide exactly:
100 units of capacity?
Not necessarily.
If demand varies, this may produce frequent shortages.
The correct capacity depends on:
- variance,
- reliability target,
- cost of spare capacity,
- cost of shortage.
So:
EXPECTED DEMAND
is only one input to:
RESOURCE DESIGN.
Overbooking is an expected-value problem
Airlines know that some passengers will not arrive.
Suppose a plane has:
100 seats.
The airline might sell:
105 tickets
because expected attendance is below 100.
This can improve capacity utilisation.
But occasionally all 105 passengers arrive.
Then the service must decide:
who receives the scarce seats?
Expected-value optimisation has created an allocation problem.
Expected-value efficiency can create tail failures
Many systems can improve average utilisation by operating closer to capacity.
But doing so leaves less protection against unexpected demand.
For example:
more efficient average utilisation
may mean:
greater risk of shortage during peaks.
This creates a fundamental trade-off between:
efficiency
and:
reliability.
We will return to this in service design.
Averages work well until the tails matter
Suppose normal conditions dominate:
99.9% of the time.
An expected-value design may work beautifully.
But if the remaining:
0.1%
contains catastrophic failures, those events may dominate system design.
Engineering often focuses heavily on:
what happens when the average is wrong.
Expected value and fairness
Suppose two resource-allocation mechanisms produce the same total expected benefit.
Mechanism A
Benefits are distributed evenly.
Mechanism B
One group receives almost everything.
From an aggregate expected-value perspective, they may appear equivalent.
From a fairness perspective, they are not.
So social systems may care about:
distribution of outcomes across people
as well as:
distribution of uncertain outcomes through probability.
These are two different meanings of "distribution".
Both matter.
Who bears the downside?
Suppose a policy produces positive expected value overall.
But:
- benefits go to one group,
- risks fall on another.
The aggregate may look attractive.
Yet the design may be unfair.
A good decision framework therefore asks:
- expected value for whom?
- expected cost for whom?
- who experiences the tail risks?
Probability-weighted does not mean morally weighted
Expected value weights outcomes by:
probability.
It does not automatically weight them by:
- fairness,
- rights,
- vulnerability,
- moral importance.
Those considerations have to enter elsewhere through:
- objectives,
- constraints,
- utility functions,
- service rules.
Mathematics cannot decide them automatically.
A simple decision framework
We can now describe a basic uncertain decision problem.
For each action:
- Identify possible outcomes.
- Assign probabilities.
- Assign values or costs to outcomes.
- Calculate expected value.
- Examine spread and tail risks.
- Check constraints.
- Consider who experiences the outcomes.
Expected value is therefore a starting point.
Not the whole decision.
Expected value in the course framework
We can extend the course's core loop:
PAST DATA
↓
MODEL
↓
PROBABILITY DISTRIBUTION OVER FUTURES
↓
EXPECTED VALUES
↓
OBJECTIVES + CONSTRAINTS
↓
DECISION
↓
ACTION
↓
FUTURE
Expected value compresses the distribution into something useful for decision-making.
But the distribution should not disappear.
The central idea
Expected value is the probability-weighted average of possible outcomes.
It gives us a powerful way to summarise uncertainty and compare choices.
But:
The expected outcome is not necessarily the outcome we will experience.
Two situations can have the same expected value and radically different:
- risks,
- uncertainty,
- tail events,
- distributional consequences.
So a good decision system should not ask only:
What is the expected value?
It should also ask:
How uncertain is it?
What happens in the tails?
Who experiences the downside?
Are there constraints we refuse to violate?
This leads directly to the next lesson.
Expected value tells us where the centre of a distribution lies.
Now we need to understand how widely outcomes are spread around that centre.
That is the role of variance and uncertainty.