Module 2 — Data: turning the world into information
Lesson 10 of 13
Sampling and selection
We rarely observe everything.
A survey does not ask every person in a country.
A medical study does not include every patient.
A traffic sensor does not record every road.
A machine-learning dataset does not contain every possible example.
Instead, we observe a sample.
A sample is a subset of a larger population, process or set of possible observations.
That creates an immediate question:
Does the data we collected represent the thing we are trying to understand?
This is the central problem of sampling and selection.
A dataset can be enormous and still be misleading if the observations it contains are not representative of the wider system.
Population and sample
Suppose we want to understand commuting behaviour across a city.
The population might be:
everyone who travels within the city
But perhaps we collect data from:
10,000 users of one transport app
Those 10,000 people form our sample.
We then hope that patterns in the sample tell us something useful about the wider population.
Conceptually:
POPULATION
↓
SELECTION PROCESS
↓
SAMPLE
↓
ANALYSIS
↓
INFERENCE ABOUT POPULATION
The quality of that inference depends heavily on the selection process.
We usually cannot measure everyone
Why not simply observe the whole population?
Sometimes we can.
A digital service may record every transaction made through its own system.
But often complete observation is:
- too expensive,
- too slow,
- physically impossible,
- unnecessary.
Imagine asking every person in a country about:
- income,
- health,
- voting intention,
- transport behaviour.
Or installing sensors on every:
- road,
- building,
- machine,
- river,
- electricity connection.
Sampling allows us to learn from a manageable subset.
The challenge is choosing that subset well.
A small good sample can beat a huge bad sample
Suppose we want to estimate public transport use.
Dataset A contains:
2,000 carefully sampled residents
Dataset B contains:
10 million journeys recorded by one smartphone app
Which is better?
It depends.
Dataset B is much larger.
But perhaps the app is used primarily by:
- younger people,
- commuters,
- people with newer smartphones,
- residents of cities.
It may poorly represent:
- older people,
- rural residents,
- people without smartphones.
A carefully designed smaller sample may provide a better estimate of the population.
More data is not automatically more representative data.
Random sampling
One ideal approach is random sampling.
Every member of the population has some known chance of being selected.
Suppose a school has:
1,000 students
and we randomly choose:
100.
If the selection is genuinely random, the sample has a reasonable chance of resembling the wider school.
Random sampling helps prevent the researcher from deliberately or accidentally choosing only particular kinds of observations.
But randomness does not guarantee perfection.
A random sample can still differ from the population by chance.
Sampling variation
Imagine repeatedly selecting random samples of 100 people from the same population.
Each sample will be slightly different.
One may contain more younger people.
Another more older people.
One may produce:
52% support
for something.
Another:
48%.
The population did not change.
The sample did.
This natural variation is called sampling variation.
It is one reason estimates from samples come with uncertainty.
Larger samples usually reduce sampling uncertainty
Suppose we estimate average electricity consumption using:
10 households
Our estimate may vary substantially depending on which ten households we happened to choose.
Use:
10,000 households
and random differences are more likely to average out.
Larger samples can therefore reduce uncertainty caused by random sampling.
But this comes with an important warning:
A larger sample reduces random sampling error. It does not automatically remove systematic selection bias.
A billion observations from the wrong population can still give the wrong answer.
Selection bias
Selection bias occurs when the process determining which observations enter the dataset systematically favours some parts of the population over others.
Suppose we ask:
How satisfied are people with the railway?
but conduct the survey only:
on trains.
Who is missing?
People who stopped using the railway because they were dissatisfied.
The sample may systematically overrepresent current users.
The resulting satisfaction estimate may therefore be biased.
Convenience sampling
Sometimes we collect the data that is easiest to obtain.
This is called convenience sampling.
For example:
- interviewing people outside one university,
- analysing users of one website,
- using patients from one hospital,
- collecting images from one online platform.
Convenience samples are often useful.
Many real datasets begin this way.
But we should be careful about generalising beyond the population they actually represent.
Internet data is a sample too
Modern AI systems are often trained on enormous amounts of internet data.
It can be tempting to think:
The dataset is so large that it must represent humanity.
But the internet itself is a selection mechanism.
To appear in an online dataset, information often had to be:
- digitised,
- published,
- accessible,
- written in a form that could be collected.
Some languages are far more represented online than others.
Some communities publish far more digital content.
Some historical knowledge was never digitised.
Some information is private.
Some content has been deleted or restricted.
So:
THE INTERNET ≠ ALL HUMAN KNOWLEDGE
It is an enormous but selective sample of human cultural production.
Who produces the data?
This is a powerful question.
Suppose we train a language model on public text.
Who creates large quantities of public text?
Perhaps:
- journalists,
- academics,
- businesses,
- governments,
- bloggers,
- software developers,
- social-media users.
Different groups contribute different amounts.
The resulting dataset may therefore reflect the voices of some people much more strongly than others.
The model learns from what was available.
Availability is itself the result of social and technological systems.
Who gets measured?
Consider healthcare data.
We might have detailed records of people who visit hospitals.
But what about people who:
- cannot access healthcare,
- avoid doctors,
- live far from facilities,
- cannot afford treatment?
Their health conditions may be underrepresented in the data.
A dataset of healthcare use is therefore not necessarily the same as a dataset of healthcare need.
This distinction can be crucial.
Observed demand is not always underlying demand
Suppose a clinic observes:
100 appointments per day
Does that mean local demand for healthcare is exactly 100 appointments?
Not necessarily.
Perhaps the clinic can only provide 100 appointments.
There might be:
150 people wanting appointments
but only 100 receive them.
The observed data reflects both:
demand
and:
the capacity of the service.
This is an important selection effect.
What we observe can be constrained by the system that produced the observation.
Housing provides another example
Suppose we want to estimate the demand for affordable housing.
We look at:
completed applications
But perhaps the application process is:
- complicated,
- slow,
- difficult to access.
Some people who need housing may never apply.
The observed applicants are a selected subset of people with underlying need.
Once again:
OBSERVED BEHAVIOUR ≠ COMPLETE UNDERLYING POPULATION
Self-selection
Sometimes people choose whether to enter the sample.
This is self-selection.
Suppose an online newspaper asks:
Do you support this policy? Vote now.
Who responds?
Often people with particularly strong opinions.
Someone who feels indifferent may not bother.
The resulting poll may therefore exaggerate extreme views.
The selection process is related to the variable being measured.
Voluntary surveys
Customer reviews have a similar problem.
Who leaves reviews?
Perhaps people who are:
- extremely happy,
- extremely unhappy.
Customers with ordinary experiences may remain silent.
A product rating might therefore reflect the experiences of people motivated to respond rather than every customer.
That does not make reviews useless.
It means we need to understand the sample.
Survivorship bias
Selection can also happen because some observations disappear before we analyse them.
Imagine studying successful companies.
We examine the strategies of firms that survived for 20 years.
We may conclude:
Successful companies tend to follow these practices.
But what about companies that followed the same practices and failed?
If they are absent from the dataset, we may draw the wrong conclusion.
This is survivorship bias.
We observe the survivors.
The failures are missing.
The classic aircraft example
A famous illustration involves aircraft returning from combat.
Engineers examine where returning aircraft have been damaged.
They see many bullet holes in particular parts of the aircraft.
It may seem logical to reinforce those areas.
But there is a problem.
The dataset contains only:
aircraft that returned.
Aircraft hit in other critical locations may not have returned at all.
The absence of damage in those areas among surviving aircraft is therefore precisely what makes those areas important.
This illustrates a deeper principle:
Sometimes the observations that are absent contain the most important information.
Selection can happen after a decision
Suppose a bank wants to predict whether applicants will repay loans.
Historical data contains outcomes for people who:
received loans.
But there is no repayment outcome for people who were rejected.
We do not know whether they would have repaid.
So the training dataset contains a selected population created by previous lending decisions.
Conceptually:
ALL APPLICANTS
↓
OLD APPROVAL POLICY
↓
APPROVED APPLICANTS
↓
OBSERVED REPAYMENT OUTCOMES
The old decision system helped determine the data available to the new model.
Feedback creates selection
This becomes particularly important when AI systems are deployed.
Suppose an AI predicts that some job applicants are unlikely to succeed.
Those applicants are not interviewed.
Because they are not hired, the organisation never observes how well they would have performed.
The model's decisions affect which future labels become observable.
The loop becomes:
MODEL
↓
SELECTION
↓
OBSERVED OUTCOMES
↓
NEW TRAINING DATA
↓
MODEL
This can reinforce existing patterns.
Recommendation systems select what users see
Suppose an online platform predicts which content a user will enjoy.
It shows:
Content A
but not:
Content B.
The user clicks Content A.
The platform records:
user likes Content A
But did the user dislike Content B?
We do not know.
They never saw it.
The system selected the opportunities from which future behaviour could be observed.
This is another powerful form of selection.
Exposure matters
Many datasets record outcomes only after someone has been exposed to an option.
We observe:
clicked / not clicked
for items shown to a user.
But items that were never shown have no click outcome.
We observe:
purchased / not purchased
for products the customer encountered.
But perhaps they would have purchased something they never discovered.
This creates a distinction between:
preference
and:
observed behaviour under a particular set of opportunities.
Sampling can be spatial
Suppose we install air-quality sensors across a country.
Where should they go?
If most sensors are placed in cities, rural areas are underrepresented.
If most sensors are placed beside roads, average pollution may appear higher than if sensors were distributed uniformly.
The sensor-placement strategy creates a spatial sample.
The resulting map depends partly on where we chose to observe.
Sampling can be temporal
Suppose we measure electricity demand only during:
09:00–17:00.
We may completely miss:
- morning peaks,
- evening peaks,
- overnight behaviour.
The observations are selected through time.
Likewise, a wildlife survey carried out only during summer may not represent winter behaviour.
Sampling must therefore consider:
when
as well as:
who and where.
Sampling frequency is also selection
A sensor recording once every hour observes a different dataset from one recording every second.
The physical process may be the same.
But we selected a different set of moments to observe.
This means the sampling frequency determines which events become visible.
A ten-second anomaly may disappear completely from hourly data.
Stratified sampling
Sometimes we know the population contains important groups.
Suppose a country contains:
- urban residents,
- rural residents.
A purely random sample might, by chance, include very few rural participants.
We could instead deliberately sample from each group.
This is called stratified sampling.
For example:
sample urban population
and:
sample rural population
separately.
This ensures both groups are represented.
Representation may require deliberate oversampling
Suppose a medical condition is rare.
Only:
1 in 10,000
people have it.
A random sample of 1,000 people may contain no positive examples at all.
That would make it difficult to train a model.
We might deliberately collect more examples of people with the condition.
The training sample no longer matches the natural population frequency.
That is not necessarily wrong.
It may be essential for learning.
But we need to remember that the training distribution was deliberately altered.
Balanced datasets can be useful but artificial
Suppose real fraud occurs in:
0.1% of transactions.
A training dataset containing:
50% fraud
and:
50% legitimate transactions
may make it easier to learn patterns.
But the model is now being trained in a world where fraud appears far more common than it is in reality.
If that difference is not handled correctly, predicted probabilities can become misleading.
Again:
The training sample does not have to perfectly match reality, but we need to understand how it differs.
Training data is a sample from a larger world
Machine learning always involves some form of sampling.
A model might see:
one million images
But there are vastly more possible images it could encounter.
It might see:
a trillion words
But there are endlessly many possible sentences.
The training dataset is a sample from a much larger space of possible experiences.
The model must then generalise beyond the examples it has seen.
That is one of the central challenges of machine learning.
The deployment population may differ
Suppose a medical model is trained using patients from:
Hospital A.
It performs extremely well.
Then it is deployed at:
Hospital B.
But Hospital B serves a different population.
Perhaps:
- ages differ,
- disease prevalence differs,
- measurement equipment differs,
- treatment practices differ.
The training sample no longer matches the deployment environment.
This can reduce performance.
Selection creates distribution shift
Suppose a model was trained using:
yesterday's users
but the service becomes much more popular.
New users are:
- younger,
- from different countries,
- using different devices.
The population has changed.
Even if the model learned the historical sample perfectly, the deployment data may now come from a different distribution.
This is one form of distribution shift.
Historical datasets select a time period too
Suppose we train an employment model using records from:
2000–2020.
That period itself is a sample of history.
It contains particular:
- economic conditions,
- technologies,
- institutions,
- social norms.
The future may not behave like that period.
So historical selection matters.
The question is not merely:
How much history do we have?
It is:
Which historical world does our dataset represent?
Sampling and fairness
Selection becomes especially important when different groups are represented unequally.
Suppose an image-recognition dataset contains:
many examples from Group A
and:
few examples from Group B.
The model has more opportunities to learn variation within Group A.
Performance may therefore differ between groups.
This is one way bias can enter an AI system.
The issue may begin before:
- architecture,
- training,
- optimisation.
It begins with who entered the dataset.
Equal sample sizes are not always the answer
Fair representation does not necessarily mean every group needs exactly the same number of examples.
Suppose one group contains far more variation in a feature relevant to the task.
Or some situations are much harder to predict.
The appropriate sampling strategy depends on:
- the objective,
- the population,
- the consequences of error.
Fair dataset design requires understanding the problem rather than mechanically equalising counts.
Labels can also be selectively observed
Consider medical diagnosis.
Some patients receive advanced diagnostic tests.
Others do not.
So high-quality labels may exist only for a selected subset of patients.
A model trained on those labels may therefore learn from a population that differs from the wider population.
This is sometimes easy to overlook.
Not only can features be selected.
Labels can be selected too.
Selection can be caused by cost
Suppose an expensive test provides an excellent measure of disease.
But only high-risk patients receive it.
We therefore have:
good labels for high-risk patients
and:
weak or missing labels for low-risk patients.
The resulting training data reflects the economics of the healthcare system.
The dataset is shaped by resource allocation.
This connects directly to the later part of the course.
Measurement resources determine the sample
Suppose an environmental agency has enough funding for:
100 sensors.
It must choose where to place them.
Perhaps it prioritises:
- cities,
- industrial zones,
- known pollution hotspots.
That may be entirely sensible.
But the resulting dataset is deliberately selected.
It should not automatically be treated as though every location had an equal chance of being observed.
Measurement itself is a resource-allocation process.
Selection can be intelligent
Not all selection is a problem.
Sometimes we deliberately choose the most informative observations.
Suppose a machine-learning model is uncertain about a particular kind of example.
Instead of randomly labelling more data, we ask humans to label examples where the model is most uncertain.
This is one form of active learning.
The selection is deliberately non-random.
The goal is to gain maximum useful information from limited labelling resources.
Active sampling
A robot may do something similar.
Suppose it is mapping a building.
It already understands one corridor extremely well.
Another area is highly uncertain.
The robot can choose to explore the uncertain area.
So:
CURRENT KNOWLEDGE
↓
IDENTIFY INFORMATION GAP
↓
SELECT NEXT OBSERVATION
↓
UPDATE KNOWLEDGE
Selection becomes part of intelligent behaviour.
Random is not always optimal
Random sampling is valuable because it supports unbiased inference about populations.
But if the goal is:
learn as efficiently as possible
rather than:
estimate a population statistic
we may prefer targeted observations.
For example, if a model already understands common examples well, additional common examples may provide little new information.
Rare or uncertain examples may be more valuable.
The right sampling strategy depends on the goal.
Exploration versus exploitation
This leads to an important problem that will appear again later.
Suppose a recommendation system knows that User A usually likes comedy.
It can:
exploit
that knowledge by recommending more comedy.
Or it can:
explore
by occasionally showing something different.
If it always exploits, it may never discover that the user also loves documentaries.
Selection therefore affects what the system can learn.
This is the exploration versus exploitation problem.
Selection shapes future knowledge
Imagine a system always recommends what it already believes users prefer.
Users mostly interact with those recommendations.
The dataset increasingly contains evidence supporting the existing belief.
Alternative preferences remain poorly observed.
The loop becomes:
CURRENT BELIEF
↓
SELECT CONTENT
↓
OBSERVE RESPONSE
↓
DATA REINFORCES BELIEF
This can create narrow feedback loops.
Exploration deliberately breaks that loop by collecting information from less certain alternatives.
Search results create selection too
Suppose a search engine ranks ten million possible pages.
A user sees only the first ten.
Clicks are therefore observed mainly among highly ranked pages.
Pages ranked low receive little opportunity to generate evidence of relevance.
If click data is later used to improve ranking, the system learns partly from outcomes produced by its own previous ranking decisions.
This is a selection feedback loop.
Observational data differs from experiments
Suppose we observe that people taking Medication A recover more frequently.
Can we conclude:
Medication A causes recovery?
Not necessarily.
Perhaps doctors tend to give Medication A to patients who were more likely to recover anyway.
The treatment groups were selected rather than randomly assigned.
This is one reason randomised experiments are so powerful.
Random assignment can help separate:
selection effects
from:
treatment effects.
We will return to this when we study correlation and causation.
Sampling defines what we can generalise to
Suppose we conduct a study using:
university students aged 18–22.
The conclusions may be valid for those students.
But can we automatically generalise them to:
- children,
- older adults,
- different countries?
Not necessarily.
The appropriate population for inference depends on the sample.
This is called external validity.
A model or study may perform excellently within its sample while generalising poorly elsewhere.
The target population
Before collecting data, we should ask:
Who or what do we ultimately want to make claims about?
That is the target population.
Then ask:
Who or what can actually enter the dataset?
That is closer to the sampling frame.
The two may differ.
For example:
Target population → all adults
Sampling frame → adults with registered telephone numbers
People outside the sampling frame cannot be selected.
That creates a coverage gap.
Coverage bias
Suppose a survey uses only an online form.
People without reliable internet access are less likely to participate.
The sampling method has created coverage bias.
The problem is not that the collected answers are false.
The issue is that some parts of the population had less opportunity to appear in the dataset.
This is exactly why we should ask:
Who had the opportunity to become data?
Non-response
Even if someone is selected, they may not respond.
Suppose 10,000 people receive a survey.
Only 2,000 reply.
If respondents and non-respondents are similar, perhaps the problem is limited.
But if people with particular views are much more likely to respond, the results may be biased.
The final dataset is therefore shaped by two selections:
who was invited
and:
who responded.
Sampling weights
Sometimes we know that groups have been sampled at different rates.
We can use weights to adjust estimates.
Suppose rural residents represent:
20% of the population
but only:
10% of the sample.
We may give rural observations more weight when estimating population-level statistics.
Weighting does not solve every sampling problem.
But it can help correct known differences between the sample and population.
Reweighting in machine learning
Similar ideas appear in machine learning.
Suppose some training examples are overrepresented.
We may assign different weights to observations during training.
The model then pays more attention to certain examples.
This can help address:
- class imbalance,
- underrepresented groups,
- different error costs.
Again, the dataset does not simply determine the learning process.
We can decide how observations should influence it.
Sampling and rare events
Rare events create a special challenge.
Suppose equipment failure occurs only:
once in every 100,000 operating hours.
A random dataset may contain very few failures.
Yet those failures may be exactly what we care about predicting.
We may therefore deliberately collect:
- failure cases,
- near-failures,
- unusual conditions.
Rare events often require targeted sampling.
AI often learns from what is available, not what is ideal
Real datasets are frequently created for reasons other than machine learning.
A hospital database exists to operate a hospital.
A financial database exists to process transactions.
A social network records interactions to run the service.
Later, someone decides to train an AI model using those records.
The dataset was not necessarily designed as a representative scientific sample.
Understanding the original data-generating process becomes essential.
Administrative data is selected by institutions
Suppose we analyse crime records.
Those records reflect:
- actual events,
- which events were reported,
- policing activity,
- enforcement priorities,
- recording practices.
Recorded crime is therefore not identical to:
all crime that occurred.
The measurement and selection systems matter.
A predictive model trained on the records learns patterns in the recorded system.
Data generated by services reflects service design
Suppose a public service offers three appointment slots per hour.
Historical data shows:
three appointments per hour.
Does that reveal the true demand?
No.
The observed data has been capped by service capacity.
Similarly:
observed queue length
may depend on:
- how long people are willing to wait,
- whether they abandon the queue,
- whether they know the service exists.
Data generated by services often contains the imprint of the service itself.
Selection and the recurring feedback loop
Recall our course framework:
PAST → DATA → MODEL → PREDICTION → DECISION → ACTION → FUTURE
Sampling and selection sit inside the transition from:
PAST → DATA.
But once models affect decisions, selection also appears later:
MODEL
↓
DECISION
↓
WHO GETS AN OPPORTUNITY
↓
WHO GENERATES AN OUTCOME
↓
NEW DATA
The dataset used tomorrow may therefore be partly selected by the model deployed today.
AI can create its own future sampling bias
Suppose a predictive-policing system sends more patrols to areas predicted to have higher crime.
More patrols can detect more incidents.
Those incidents enter the dataset.
The area appears to have even more recorded crime.
The model may send even more patrols.
The loop becomes:
PREDICTION
↓
RESOURCE ALLOCATION
↓
MORE OBSERVATION
↓
MORE RECORDED EVENTS
↓
STRONGER PREDICTION
Whether or not the original prediction was perfect, the measurement process has changed.
This is a classic example of why prediction systems cannot be understood separately from the systems in which they operate.
Selection affects labels as well as features
Suppose we train a model to predict:
employee performance
But historical performance scores exist only for employees who were hired.
Applicants who were rejected have no performance label.
So:
HIRING DECISION
determined:
WHO COULD GENERATE A PERFORMANCE OUTCOME.
The model may therefore learn within the world created by previous selection decisions.
This problem appears in:
- hiring,
- credit,
- education,
- healthcare,
- insurance.
We often cannot observe counterfactuals
Suppose a university rejects an applicant.
Would that applicant have succeeded if admitted?
We do not know.
Suppose a bank rejects a loan.
Would the borrower have repaid?
We do not know.
Suppose a patient receives Treatment A.
What would have happened under Treatment B?
We usually cannot observe both outcomes for the same person.
The unobserved alternative is a counterfactual.
Selection determines which outcome becomes visible.
This is fundamental to causal inference.
Selection is not necessarily bad
This deserves emphasis.
Selection is unavoidable.
We cannot observe:
- every person,
- every place,
- every time,
- every possible future.
We must choose where to collect information.
The goal is not to eliminate selection.
It is to:
- understand it,
- design it deliberately,
- account for its effects,
- represent the uncertainty it creates.
This is similar to missing data.
The existence of selection does not mean the dataset is unusable.
It means we need to reason carefully about what the dataset represents.
We do not need a perfect sample for every task
Suppose our goal is to detect a particular mechanical fault.
We may not care whether the dataset represents the average machine in the world.
We care whether it contains enough information to distinguish:
fault
from:
normal operation.
For another task, such as estimating national failure rates, representativeness may be essential.
Sampling quality is therefore always relative to the question.
Prediction and population estimation are different goals
Suppose a dataset overrepresents young people.
That could be a major problem if we want to estimate:
average national income.
But perhaps it is less problematic if the model will only ever be deployed to:
young users of the same service.
We need to distinguish between:
What population produced the data?
and:
What population will receive the prediction?
The closer those match, the more confident we may be in generalisation.
Selection should match deployment
Suppose an autonomous vehicle is trained mostly using:
- sunny weather,
- daytime driving,
- wide roads.
Then it is deployed:
- at night,
- in snow,
- on narrow rural roads.
The deployment environment contains situations poorly represented in the sample.
The problem is not necessarily the model architecture.
The training experience did not sufficiently represent the world in which the model was later asked to operate.
We can deliberately collect edge cases
An intelligent dataset does not need to mirror frequency perfectly.
Suppose dangerous situations are rare but extremely important.
We may deliberately collect more examples of:
- pedestrians entering roads,
- unusual vehicle types,
- emergency vehicles,
- extreme weather.
These are sometimes called edge cases.
Their real-world frequency may be low.
Their importance may be high.
Data collection should therefore consider consequences, not merely frequency.
Sampling is a resource-allocation problem
Collecting data costs resources.
We have finite:
- sensors,
- human labellers,
- storage,
- money,
- time,
- compute.
So we need to decide:
Which observations are worth collecting?
This is itself an optimisation problem.
Perhaps we should collect more data:
- where uncertainty is high,
- where errors are costly,
- where populations are underrepresented,
- where the system is changing quickly.
The design of the dataset can therefore become intelligent.
The value of an observation
Suppose a model already has millions of nearly identical examples.
One more may add almost no information.
Another example from a rare condition may dramatically improve understanding.
The value of data therefore depends on:
- what we already know,
- what we remain uncertain about,
- what decisions the model supports.
A useful question is:
What observation would most improve our ability to make the next decision?
From passive datasets to active learning
Traditional datasets are often treated as fixed.
We collect data.
Then train the model.
But intelligent systems can create a loop:
MODEL
↓
IDENTIFY UNCERTAINTY
↓
SELECT NEW DATA TO COLLECT
↓
UPDATE MODEL
This is active learning.
The model participates in deciding what it needs to learn next.
That is a powerful extension of the course's feedback framework.
Selection and exploration
Suppose an AI knows a great deal about:
Situation A
but very little about:
Situation B.
If it always chooses actions associated with Situation A, it may never learn much about B.
Exploration deliberately gathers information from less familiar parts of the state or action space.
This becomes particularly important in:
- reinforcement learning,
- recommendation systems,
- robotics,
- adaptive services.
Learning depends on what the system chooses to experience.
Ask who is missing
Whenever you encounter a dataset, ask:
- What is the target population?
- What population actually produced the data?
- How were observations selected?
- Who had the opportunity to be included?
- Who did not?
- Was participation voluntary?
- Did some groups respond more than others?
- Did geography affect inclusion?
- Did time affect inclusion?
- Were rare cases deliberately oversampled?
- Did previous decisions determine who generated labels?
- Is the deployment population different from the training population?
- Could the system's own predictions change what gets observed next?
One of the most powerful questions in data analysis is simply:
Who or what is not in this dataset?
The central idea
A dataset is not simply a collection of facts.
It is a collection of selected observations.
Between the world and the dataset sits a selection process:
WORLD
↓
WHO / WHAT / WHEN / WHERE DO WE OBSERVE?
↓
SAMPLE
↓
DATASET
↓
MODEL
That selection process shapes what the model can learn.
A huge dataset can still be unrepresentative.
A small dataset can still be highly informative.
A deliberately biased sample can sometimes be useful for learning rare events.
The important thing is to understand what population the data represents and what inference we are trying to make.
We do not need to observe everything. We need to understand how what we observed relates to what we did not observe.
And once we recognise that datasets are samples rather than complete representations of reality, uncertainty becomes unavoidable.
But even the observations that do make it into the sample are not perfect.
Sensors have tolerances.
People make mistakes.
Measurements contain noise.
That leads to the next lesson:
measurement error — why an observation can be useful without being exact, and how intelligent systems can reason when the numbers themselves are uncertain.