Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

2.2 Systematic and Statistical Uncertainty: Difference between revisions

From Sense & Sensibility & Science
// Edit via Wikitext Extension for VSCode
Tag: Manual revert
// via Wikitext Extension for VSCode
 
(15 intermediate revisions by 2 users not shown)
Line 1: Line 1:
[[File:Topic Cover - 2.2 Systematic and Statistical Uncertainty.png|thumb]]
{{Cover|2.2 Systematic and Statistical Uncertainty}}


This topic explores the sources of error and uncertainty in data.
Any measurement by an instrument comes with its inevitable imperfections. How do we quantify and communicate the extent to which the reading on an instrument can be trusted? We classify reasons why the reading may differ from the true value into two broad categories—systematic uncertainty and statistical uncertainty. We must learn to deal with uncertainty in our knowledge.
 
{{Navbox}}


== The Lesson in Context ==
== The Lesson in Context ==


<!-- Always begin section with a description of this lesson in relation to the course as a whole. -->
<!-- Always begin section with a description of this lesson in relation to the course as a whole. -->
This lesson teaches students that the inevitable imperfections of instrumental measurements of the real world can be quantified and studied in their own rights. They are categorized into statistical uncertainty and systematic uncertainty. We will teach them to be aware of the sources of uncertainty in each measurement and some elementary ways to mitigate them. This is illustrated by the human histogram activity, in which students can see statistical distributions and physically experience the effects of systematic uncertainty. We will also discuss how systematic uncertainties affected results of political polling in the 2016 US presidential election.
We physically illustrate the difference between statistical and systematic uncertainty with the human histogram activity, in which every student gets to participate as a data point. We will also discuss how systematic uncertainties affected results of political polling in the 2016 US presidential election. This is the first in a series of lessons that familiarize students with the important concept of epistemic uncertainty.


<!-- Expandable section relating this lesson to earlier lessons. -->
<!-- Expandable section relating this lesson to other lessons. -->
{{Expand|Relation to Earlier Lessons|
{{Expand|Relation to Other Lessons|
'''Earlier Lessons'''
{{ContextLesson|1.2 Shared Reality and Modeling}}
{{ContextLesson|1.2 Shared Reality and Modeling}}
{{ContextRelation|It is inevitable that our experience or measurement of the external reality is imperfect. This lesson's concepts help to quantify these imperfections.}}
{{ContextRelation|It is inevitable that our experience or measurement of the external reality is imperfect. This lesson's concepts help to quantify these imperfections.}}
{{ContextLesson|2.1 Senses and Instrumentation}}
{{ContextLesson|2.1 Senses and Instrumentation}}
{{ContextRelation|No instrument is perfect. Systematic and statistical uncertainties help quantify these imperfections and allow us to compare two different instruments or methods of measurement.}}
{{ContextRelation|No instrument is perfect. Systematic and statistical uncertainties help quantify these imperfections and allow us to compare two different instruments or methods of measurement.}}
}}
{{Line}}
<!-- Expandable section relating this lesson to later lessons. -->
'''Later Lessons'''
{{Expand|Relation to Later Lessons|
{{ContextLesson|3.1 Probabilistic Reasoning}}
{{ContextRelation|Instrumental uncertainty can be expressed as error bars and confidence intervals. These translate to a probabilistic understanding of where the true value lies.}}
{{ContextLesson|6.1 Correlation and Causation}}
{{ContextLesson|6.1 Correlation and Causation}}
{{ContextRelation|Randomized assignment is one way to remove the systematic uncertainty by making sure that the intervention and control groups are not correlated with some other variable related to the method of assignment itself, e.g. a male vs. female group in a drug trial.}}
{{ContextRelation|Randomized assignment is one way to remove the systematic uncertainty by making sure that the intervention and control groups are not correlated with some other variable related to the method of assignment itself, e.g. a male vs. female group in a drug trial.}}
{{ContextRelation|Placebo effect is a systematic uncertainty in the measurement of the effectiveness of a treatment. Therefore, we must "subtract" the effect of the placebo treatment from the effect of the real treatment.}}
{{ContextRelation|Placebo effect is a systematic uncertainty in the measurement of the effectiveness of a treatment. Therefore, we must "subtract" the effect of the placebo treatment from the effect of the real treatment.}}
}}
}}
== Takeaways ==
== Takeaways ==


Line 50: Line 49:
{{BoxCaution|Because a proxy is not a direct measure of the quantity of interest, it is a place that systematic bias can creep in. For example, using self-report ratings of happiness as a measure of happiness may be subject to cultural differences and/or comparison effects.}}
{{BoxCaution|Because a proxy is not a direct measure of the quantity of interest, it is a place that systematic bias can creep in. For example, using self-report ratings of happiness as a measure of happiness may be subject to cultural differences and/or comparison effects.}}
{{Definition|Triangulation|A method of dealing with systematic errors inherent in a type of measurement by attempting to measure the same phenomenon in multiple different ways or through lots of different proxies.}}
{{Definition|Triangulation|A method of dealing with systematic errors inherent in a type of measurement by attempting to measure the same phenomenon in multiple different ways or through lots of different proxies.}}
<br />


|-|Examples=
|-|Examples=


<!-- Example formatting is still experimental. -->
{{Example
'''Proxy Example'''
|Proxy Example
: When measuring crime rate, it is only possible to measure the rate of reported crime, making it a proxy of the true crime rate. Even when measuring temperature, it is only possible to observe the reading on an instrument, also a form of proxy.
|When measuring crime rate, it is only possible to measure the rate of reported crime, making it a proxy of the true crime rate. Even when measuring temperature, it is only possible to observe the reading on an instrument, also a form of proxy.}}
{{Line}}
{{Example
'''Simple Systematic Uncertainty Example'''
|Simple Systematic Uncertainty Example
: "It won't do us any good to average lots of test subjects' heights together if our tape measure got shrunk in the wash!"
|"It won't do us any good to average lots of test subjects' heights together if our tape measure got shrunk in the wash!"}}
{{Line}}
{{Example
'''Simple Statistical Uncertainty Example'''
|Simple Statistical Uncertainty Example
: "Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. The people they are talking to might not be representative of the whole population. For example, older people can be more likely to pick up the phone and talk to pollsters, so there might be a systematic bias in that direction."  
|"Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. The people they are talking to might not be representative of the whole population. For example, older people can be more likely to pick up the phone and talk to pollsters, so there might be a systematic bias in that direction."}}
{{Line}}
{{Example
'''Polling Models'''
|Polling Models
: "Indeed, [the 2016] election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for [systematic uncertainties]. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out." [https://www.forbes.com/sites/startswithabang/2016/11/09/the-science-of-error-how-polling-botched-the-2016-election/#accdd8c37959 (Source)]
|"Indeed, [the 2016] election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for [systematic uncertainties]. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out."
{{Line}}
|links={{LinkCard
'''Stiffness of Springs'''
|url=https://www.forbes.com/sites/startswithabang/2016/11/09/the-science-of-error-how-polling-botched-the-2016-election/
: "The stiffness of many springs depends on their temperature. If you measure the stiffness of a spring many times, by compressing and decompressing it, the internal friction inside the spring may cause it to warm. You may see this by a systematic trend in your data set; for example, each data point in a data set will be smaller than the previous one." [https://drive.google.com/open?id=1_-lEfKy-qkq9imE2MsLNOl2Asja3JxFa (Source)]
|title=The Science of Error: How Polling Botched the 2016 Election
{{Line}}
|description=Forbes article on systematic error in 2016 election polling.}}
'''Systematic Error in Life Sciences'''
}}
: "Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error — cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture." This is a more complex life science example of systematic error.
{{Example
{{Line}}
|Stiffness of Springs
'''Reducing Systematic Error'''
|"The stiffness of many springs depends on their temperature. If you measure the stiffness of a spring many times, by compressing and decompressing it, the internal friction inside the spring may cause it to warm. You may see this by a systematic trend in your data set; for example, each data point in a data set will be smaller than the previous one."
: "If we estimate the effect of a drug on weight by randomly assigning people to take the drug vs. not take it and then measure their weight after a year, we could subtract the average weight loss of drug-takers vs. non-drug-takers to get the effect size of the drug on weight loss. But the people know if they're taking a drug for weight loss, so there could be a placebo effect creating a systematic bias. So the better way to do the experiment is to give the control group sugar pills. Then we can be more confident that any weight loss is due to the drug, and not a systematic bias created by the placebo effect."
|links={{LinkCard
{{Line}}
|url=https://drive.google.com/open?id=1_-lEfKy-qkq9imE2MsLNOl2Asja3JxFa
'''Sleep Questionnaire'''
|title=Sources of Systematic Error
: If you asked just a handful of random people on the street how much they slept the night before, the average of their answers could be quite different from the true average of the whole population due to random differences between people (statistical uncertainty). This can be improved by asking more people (say, hundreds or thousands). However, if you asked hundreds of random people on a college campus the same question, all of their answers could be skewed in one direction due to collective sleep deprivation (systematic uncertainty), which would not be improved by asking more college students.
|description=Notes covering the spring-stiffness example.}}
{{Line}}
}}
'''Brightness of a Star'''
{{Example
: Suppose you are an astronomer measuring the brightness of a star. The star twinkles due to random atmospheric fluctuations (statistical uncertainty), but the presence of the atmosphere itself, together with clouds, always reduces the brightness of the star (systematic uncertainty).
|Systematic Error in Life Sciences
|"Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error — cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture." This is a more complex life science example of systematic error.}}
{{Example
|Reducing Systematic Error
|"If we estimate the effect of a drug on weight by randomly assigning people to take the drug vs. not take it and then measure their weight after a year, we could subtract the average weight loss of drug-takers vs. non-drug-takers to get the effect size of the drug on weight loss. But the people know if they're taking a drug for weight loss, so there could be a placebo effect creating a systematic bias. So the better way to do the experiment is to give the control group sugar pills. Then we can be more confident that any weight loss is due to the drug, and not a systematic bias created by the placebo effect."}}
{{Example
|Sleep Questionnaire
|If you asked just a handful of random people on the street how much they slept the night before, the average of their answers could be quite different from the true average of the whole population due to random differences between people (statistical uncertainty). This can be improved by asking more people (say, hundreds or thousands). However, if you asked hundreds of random people on a college campus the same question, all of their answers could be skewed in one direction due to collective sleep deprivation (systematic uncertainty), which would not be improved by asking more college students.}}
{{Example
|Brightness of a Star
|Suppose you are an astronomer measuring the brightness of a star. The star twinkles due to random atmospheric fluctuations (statistical uncertainty), but the presence of the atmosphere itself, together with clouds, always reduces the brightness of the star (systematic uncertainty).}}
{{Exemplary
|{{Blockquote|It won't do us any good to average lots and lots of test subjects' heights together if our tape measure got shrunk in the wash!}}
{{Blockquote|Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. They have no idea if the people who they are reaching in their polling might be a highly skewed segment of the public so that they have, for example, a systematic bias towards the opinions of an older than average population.}}
{{Blockquote|Indeed, this election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for them. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out.|[https://www.forbes.com/sites/startswithabang/2016/11/09/the-science-of-error-how-polling-botched-the-2016-election/ Source]}}
{{Blockquote|Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error—cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture.}}
}}


|-|Common Misconceptions=
|-|Common Misconceptions=
Line 88: Line 102:
{{Misconception|A single measurement can only have either systematic or statistical uncertainty.|Every measurement can come with systematic ''and'' statistical uncertainty, often with multiple sources to different degrees.}}
{{Misconception|A single measurement can only have either systematic or statistical uncertainty.|Every measurement can come with systematic ''and'' statistical uncertainty, often with multiple sources to different degrees.}}


</tabber>
|-|Expanded Learning Goals=
<restricted>


== Useful Resources ==
After this lesson, students should
 
# Attitudes
<tabber>
## Be aware that every measurement comes with some degree of uncertainty.
 
# Concept Acquisition
|-|Lecture Video=
## '''Statistical Uncertainty/Error:''' Differences between reality and our measurement on the basis of random imprecisions.
 
### All measurements have a certain amount of variance, which are just differences between multiple measurements due to error and/or genuine variation in the sample. These differences will not all go in the same direction.
<br /><center><youtube>d1cXjT6f7fE</youtube></center><br />
### Statistical uncertainty can be reduced by averaging a larger amount of data.
 
## '''Systematic Uncertainty/Error:''' Differences between reality and our measurement that skew our results in one direction.
|-|Bonus Video=
### Such measurements will show a consistent bias—that is, a consistent deviation from reality in one direction.
 
### Systematic uncertainty cannot be reduced by averaging a larger amount of data.
<br /><center><youtube>WT4xqZjGWUQ</youtube></center><br />
# Concept Application
 
## Identify sources of measurement uncertainty/error that introduce statistical uncertainty/error, that introduce systematic uncertainty/error, and that introduce both.
|-|Discussion Slides=
## Suggest approaches to reducing or constraining measurement uncertainty (both statistical and systematic).
 
{{LinkCard
|url=https://docs.google.com/presentation/d/1SiDji1Bf0keUJGH_KvaqP5pCrv9jWHaG9eRO3vfOUwE/
|title=Discussion Slides Template
|description=The discussion slides for this lesson.
}}
<br />
 
|-|Readings and Assignments=
 
{{LinkCardInternal
|url=:File:Why FiveThirtyEight Gave Trump a Better Chance Than Almost Anyone Else - Silver.pdf
|title=Why FiveThirtyEight Gave Trump a Better Chance Than Almost Anyone Else
|description=Reading regarding FiveThirtyEight's model for the 2016 US presidential election.}}
<br />


</tabber>
</tabber>
 
{{#restricted:{{Private:2.2 Systematic and Statistical Uncertainty}}}}
== Recommended Outline ==
{{NavCard|chapter=Lesson plans|text=All lesson plans|prev=2.1 Senses and Instrumentation|next=3.1 Probabilistic Reasoning}}
 
=== Before Class ===
 
* Prepare sheets of paper with ranges of heights for the [[#Human Histogram|Human Histogram activity]]. You may also want to scout a location and get some tape to hang the sheets with.
* Remind students to read the [[#Useful Resources|FiveThirtyEight article]].
 
=== During Class ===
 
{| class="wikitable" style="margin-left: 0px; margin-right: auto;"
|5 Minutes
|Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc.
|-
|2 Minutes
|Have the students attempt the [[#Warm-up Question|warm-up question]].
|-
|7 Minutes
|Go through the [[#Concept Review Questions|concept review questions]].
|-
|20 Minutes
|Walk the students through the [[#FiveThirtyEight Reading Discussion|FiveThirtyEight reading discussion]].
|-
|33 Minutes
|Conduct the [[#Human Histogram|Human Histogram activity]].
|-
|Remaining Time
|This leaves 13 minutes of wiggle room for any activities that go long. If you have time at the end you can collect and answer any people's lingering questions.
|}
 
== Lesson Content ==
 
=== Warm-up Question ===
 
Making lots of measurements and averaging them is a good way to reduce...
<ol style="list-style-type:lower-alpha">
<li>{{Correct|Statistical Uncertainty}}</li>
<li>Systematic Uncertainty</li>
</ol>
 
=== Concept Review Questions ===
 
<ol start=1><li>Suppose you've discovered a new substance. You've taken it into the lab and you want to see what temperature it melts at. But, you're worried that the thermometer is miscalibrated. Would this be a source of statistical or systematic uncertainty in measurements that use this thermometer? How might you figure this out and reduce this uncertainty?</li></ol>
{{BoxAnswer|Miscalibration of instrument is a source of systematic uncertainty. You can test it by putting the thermometer into just boiling water and see if it's 100 degrees Celsius, and into just freezing water and see if it's 0 degrees Celsius. OR you can compare it to another thermometer.}}
<ol start=2><li>Suppose you're running a study of the effects of class size on student learning. You find that students at a school with average class size of 40 perform worse than students at a neighboring school with average class size of 30. Both schools have 3000 students. Should you be more worried about systematic or statistical uncertainty? If you had abundant control and resources, what could you do to reduce each source of uncertainty that you can think of?</li></ol>
{{BoxAnswer|You should be more worried about systematic uncertainty because there are likely to be other differences between the two schools and the students. You can reduce systematic uncertainty by randomly assigning students to school (with bussing) and/or by randomly assigning classes to have 30 vs. 40 students, with equal numbers of small/large classrooms at each school. This takes out the confound of school (and the many variables that may differ between schools) with class size. You could reduce statistical uncertainty by doing the same thing at many schools, not only two.}}
=== Human Histogram ===
 
This activity aims to illustrate the effects of statistical and systematic uncertainties in data by having students participate as individual data points, forming a "human histogram" in which each student is actually a living data point.
<center><youtube>JqIpnp1S9NE</youtube></center>
<!-- {{LinkCard
|url=https://drive.google.com/file/d/1V91nsU86A4xVLUn8iHV_IxZ64l1KR5wV/view
|title=Live Demonstration Video
|description=A recording of the activity being done during class at UC Berkeley.}}
{{BoxWarning|Don't share the video publicly. It has student faces in it.}} -->
==== Preparation ====
 
{{BoxTip|title=Finding a Location|If possible, find a location outdoors near your classroom. This is where the activity will be conducted. So, make sure that it's a place you could quickly migrate the students to during class. The activity can be done in the classroom if you don't find a good spot. But, space may be more cramped.}}
Make sheets of paper with large print or writing to denote the bins in the histogram. If you have 30 or so students then the following four categories should be sufficient. If you have 60-80 students then it may be best to have six or so bins instead of four. Make sure you have any scotch tape to hold up the signs indoors, rocks to hold them down outside, or anything else you may need in order to have the signs marked clearly.
 
{| class="wikitable" style="margin-left: auto; margin-right: auto;"
!Imperial
!Metric
|-
|< 5'2"
|< 158 cm
|-
|5'2" - 5'7"
|158 cm - 170 cm
|-
|5'7" - 6'
|170 cm - 183 cm
|-
|> 6'
|> 183 cm
|}
 
How the activity will be carried out heavily depends on the exact venue. Make sure to familiarize yourself with the venue and rehearse it ahead of time. During or just before class, place the papers corresponding to the "bins" on opposite sides of whatever space you're conducting the activity. There should be one set of bins for football fans and one for soccer fans. Make sure there's enough space that several students can stand at each piece of paper without having the bins bleed into each other.
==== Instructions ====
 
{| class="wikitable" style="margin-left: 0px; margin-right: auto;"
|4 Minutes
|If you're not doing the activity in your classroom, take this time to move all the students outside.
|-
|13 Minutes
|Conduct the "[[#During the Game|During the Game]]" part of the activity.
|-
|16 Minutes
|Conduct the "[[#At the Party|At the Party]]" part of the activity.
|}
 
===== During the Game =====
 
{| class="wikitable" style="margin-left: 0px; margin-right: auto;"
|2 Minutes
|Split the class into two halves. Let students choose whether they'd like to be football fans or soccer fans.
|-
|3 Minutes
|Tell students that they are attending a football or soccer game as per their preference. You want to know how many people of what height are at each game. Instruct the students to walk into the bin that contains their height (e.g. for someone 5'7", go into the 5'6" - 5'9" bin) on their own side.
|-
|1 Minute
|When it settles, write down which histogram skews taller. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set.
|-
|3 Minutes
|The instructor should now role play as the "alien researcher" who is studying human heights. Eyeball the height distributions of the two groups to see which group is taller on average than the other. The alien researcher says, "it looks like the [football/soccer] fans here are taller than the [soccer/football] fans, so [football/soccer] fans must in general be taller than [soccer/football] fans!"
|-
|4 Minutes
|Ask the students the final if this alien researcher's conclusion is justified? Why or why not?
|}
{{BoxAnswer|title=Answer|The alien is not justified in their conclusion, because they made a generalization about an overall population based on only very few samples. The differences in heights that we see in this small class are subject to statistical uncertainty.}}
{{BoxTip|Some students might not have a preference between football and soccer. Try to make the two halves more or less similar in size.}}
===== At the Party =====
 
{| class="wikitable" style="margin-left: 0px; margin-right: auto;"
|3 Minutes
|A second alien researcher decides to do the same study, but at the after parties. Now tell the story that football fans are going to a beach party, where they may wear flip flops or no shoes at all. On the other hand, soccer fans are going to a fancy hat ball, where they may wear fancy hats and/or high heels.
|-
|3 Minutes
|Instruct students to now imagine how tall they may be including their new apparel (extremely flamboyant for soccer fans) and walk into the bins containing their new height. Make it clear that soccer fans should join the bin indicating their height ''including the imaginary heels and hats'', because the alien assumes the shoes and hats are part of the human organism.
|-
|1 Minute
|When it settles, write down which histogram skews taller and screenshot them/photograph. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set.
|-
|3 Minutes
|he second alien researcher says, "it looks like the soccer fans are taller than the football fans, so soccer fans must in general be taller than football fans!"
|-
|6 Minutes
|Ask the students the final if ''this'' alien researcher's conclusion is justified? Why or why not?
|}
{{BoxAnswer|title=Answer|The alien is not justified in their conclusion, not only because of statistical uncertainty, but also because there is now a major source of systematic uncertainty—the hats and shoes on the soccer fans.}}
{{BoxCaution|Conclude that statistical uncertainty could go both ways simply due to the randomness in the selection of samples (such as heights of students in this section), and it can be improved by increasing the sample size (e.g. the whole SSS class).}}
{{BoxCaution|Systematic uncertainties cause the data to be skewed in one direction, and cannot be removed just by increasing sample size. To remove this effect, the alien would have to understand human physiology—that hats and shoes are not part of the human body—and account for that.}}
{{BoxCaution|We did this activity with only the small number of students in a single discussion section. Did it work? How would it have worked better if we did it with the entire class? You can use this as a teaching opportunity to show how systematic and statistical errors manifest in small sample sizes.}}
=== FiveThirtyEight Reading Discussion ===
{{LinkCardInternal
|url=:File:Why FiveThirtyEight Gave Trump a Better Chance Than Almost Anyone Else - Silver.pdf
|title=Why FiveThirtyEight Gave Trump a Better Chance Than Almost Anyone Else
|description=Reading regarding FiveThirtyEight's model for the 2016 US presidential election.}}
Spend four minutes on each question.
<ol start=1><li>Why didn't the huge volume of polling data during the 2016 election eliminate uncertainty?</li></ol>
{{BoxAnswer|Because the polls shared systematic errors, more often reaching kinds of people who were more likely to vote for Clinton and missing the kinds of people who were more likely to vote for Trump.}}
<ol start=2><li>Why is it valuable to have multiple polls?</li></ol>
{{BoxAnswer|Each poll is a bit different, using slightly different strategies and reaching different sets of people, so they can cancel out some of each other's systematic uncertainties.}}
<ol start=3><li>Why do polls tend to replicate each other's mistakes?</li></ol>
{{BoxAnswer|Although no two polls are identical, they tend to use similar strategies for reaching likely voters and so may replicate systematic uncertainties across polls. For example, polls that depend on people picking up their phones from unknown callers will systematically miss people who don't pick up such calls, who may be systematically different. Even with absolute best practices, pollsters can only reach people who agree to answer questions, and willingness to answer a poll may be negatively correlated with antisocial tendencies or distrust in institutions, which may affect voting preferences.}}
<ol start=4><li>Why was FiveThirtyEight's prediction better than the predictions of other pundits?</li></ol>
{{BoxAnswer|FiveThirtyEight took into account the historical accuracy of the individual polls. By noting systematic discrepancies from that historical record, they were able to use more appropriate error bars taking the uncertainties into account.}}
<ol start=5><li>Why are state polls less reliable than national polls?</li></ol>
{{BoxAnswer|Because they usually have smaller sample sizes and there are fewer polls to balance out systematic errors from any given poll.}}</restricted>{{NavCard|prev=2.1 Senses and Instrumentation|next=3.1 Probabilistic Reasoning}}
[[Category:Lesson plans]]
[[Category:Lesson plans]]

Latest revision as of 22:33, 11 June 2026

Any measurement by an instrument comes with its inevitable imperfections. How do we quantify and communicate the extent to which the reading on an instrument can be trusted? We classify reasons why the reading may differ from the true value into two broad categories—systematic uncertainty and statistical uncertainty. We must learn to deal with uncertainty in our knowledge.

The Lesson in Context

We physically illustrate the difference between statistical and systematic uncertainty with the human histogram activity, in which every student gets to participate as a data point. We will also discuss how systematic uncertainties affected results of political polling in the 2016 US presidential election. This is the first in a series of lessons that familiarize students with the important concept of epistemic uncertainty.

Earlier Lessons

1.2 Shared Reality and Modeling
  • It is inevitable that our experience or measurement of the external reality is imperfect. This lesson's concepts help to quantify these imperfections.
2.1 Senses and Instrumentation
  • No instrument is perfect. Systematic and statistical uncertainties help quantify these imperfections and allow us to compare two different instruments or methods of measurement.

Later Lessons

3.1 Probabilistic Reasoning
  • Instrumental uncertainty can be expressed as error bars and confidence intervals. These translate to a probabilistic understanding of where the true value lies.
6.1 Correlation and Causation
  • Randomized assignment is one way to remove the systematic uncertainty by making sure that the intervention and control groups are not correlated with some other variable related to the method of assignment itself, e.g. a male vs. female group in a drug trial.
  • Placebo effect is a systematic uncertainty in the measurement of the effectiveness of a treatment. Therefore, we must "subtract" the effect of the placebo treatment from the effect of the real treatment.

Takeaways

After this lesson, students should

  1. Realize that our contact with reality is often mediated by measurement and quantification. We need to be aware that every measurement comes with some degree of uncertainty (deviation from the "true" value in reality).
  2. Identify sources of measurement uncertainty/error that introduce statistical uncertainty/error, that introduce systematic uncertainty/error, and that introduce both.
  3. Understand how to use repeated measures to reduce statistical uncertainty.
  4. Recognize the difficulty of removing systematic uncertainty, and that the process of science involves creativity in identifying sources of systematic uncertainty and inventing strategies to reduce or eliminate them.

Students will likely keep asking "but how can I tell between statistical and systematic uncertainties", and the answer would be to offer as many diverse examples as possible. Also if you can reduce the uncertainty simply by collecting more data, it's statistical.


Statistical Uncertainty

Differences between reality and measurements on the basis of random imprecisions or "noise."

Systematic Uncertainty

Differences between reality and measurements that skew results in one direction.

Accuracy

How close the measured value is to the true value.

Precision

How similar are all the measured values of the same thing (consistency).

A measurement can be very precise but wrong/inaccurate (low statistical uncertainty but high systematic uncertainty), or it could have a large variance between subsequent measurements but average accurately to the true value (low systematic uncertainty but high statistical uncertainty).

Proxy

An observable measurement used to approximate a quantity that is not directly observable or measurable. E.g. people's ratings of agreement with the statement "I am happy" on a scale from 1 to 7, is a measure (proxy) of their happiness, or using zip code as a proxy for socio-economic status.

Because a proxy is not a direct measure of the quantity of interest, it is a place that systematic bias can creep in. For example, using self-report ratings of happiness as a measure of happiness may be subject to cultural differences and/or comparison effects.

Triangulation

A method of dealing with systematic errors inherent in a type of measurement by attempting to measure the same phenomenon in multiple different ways or through lots of different proxies.

Proxy Example

When measuring crime rate, it is only possible to measure the rate of reported crime, making it a proxy of the true crime rate. Even when measuring temperature, it is only possible to observe the reading on an instrument, also a form of proxy.

Simple Systematic Uncertainty Example

"It won't do us any good to average lots of test subjects' heights together if our tape measure got shrunk in the wash!"

Simple Statistical Uncertainty Example

"Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. The people they are talking to might not be representative of the whole population. For example, older people can be more likely to pick up the phone and talk to pollsters, so there might be a systematic bias in that direction."

Polling Models

"Indeed, [the 2016] election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for [systematic uncertainties]. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out."

Stiffness of Springs

"The stiffness of many springs depends on their temperature. If you measure the stiffness of a spring many times, by compressing and decompressing it, the internal friction inside the spring may cause it to warm. You may see this by a systematic trend in your data set; for example, each data point in a data set will be smaller than the previous one."

Systematic Error in Life Sciences

"Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error — cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture." This is a more complex life science example of systematic error.

Reducing Systematic Error

"If we estimate the effect of a drug on weight by randomly assigning people to take the drug vs. not take it and then measure their weight after a year, we could subtract the average weight loss of drug-takers vs. non-drug-takers to get the effect size of the drug on weight loss. But the people know if they're taking a drug for weight loss, so there could be a placebo effect creating a systematic bias. So the better way to do the experiment is to give the control group sugar pills. Then we can be more confident that any weight loss is due to the drug, and not a systematic bias created by the placebo effect."

Sleep Questionnaire

If you asked just a handful of random people on the street how much they slept the night before, the average of their answers could be quite different from the true average of the whole population due to random differences between people (statistical uncertainty). This can be improved by asking more people (say, hundreds or thousands). However, if you asked hundreds of random people on a college campus the same question, all of their answers could be skewed in one direction due to collective sleep deprivation (systematic uncertainty), which would not be improved by asking more college students.

Brightness of a Star

Suppose you are an astronomer measuring the brightness of a star. The star twinkles due to random atmospheric fluctuations (statistical uncertainty), but the presence of the atmosphere itself, together with clouds, always reduces the brightness of the star (systematic uncertainty).

Exemplary Quotes

It won't do us any good to average lots and lots of test subjects' heights together if our tape measure got shrunk in the wash!

Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. They have no idea if the people who they are reaching in their polling might be a highly skewed segment of the public so that they have, for example, a systematic bias towards the opinions of an older than average population.

Indeed, this election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for them. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out.

Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error—cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture.

The words "uncertainty" and "error" mean that our instruments or measurement methods are somehow broken, deficient, or not to be trusted.

These words describe the inevitable and perfectly acceptable gap between measurement and reality.

A single measurement can only have either systematic or statistical uncertainty.

Every measurement can come with systematic and statistical uncertainty, often with multiple sources to different degrees.

After this lesson, students should

  1. Attitudes
    1. Be aware that every measurement comes with some degree of uncertainty.
  2. Concept Acquisition
    1. Statistical Uncertainty/Error: Differences between reality and our measurement on the basis of random imprecisions.
      1. All measurements have a certain amount of variance, which are just differences between multiple measurements due to error and/or genuine variation in the sample. These differences will not all go in the same direction.
      2. Statistical uncertainty can be reduced by averaging a larger amount of data.
    2. Systematic Uncertainty/Error: Differences between reality and our measurement that skew our results in one direction.
      1. Such measurements will show a consistent bias—that is, a consistent deviation from reality in one direction.
      2. Systematic uncertainty cannot be reduced by averaging a larger amount of data.
  3. Concept Application
    1. Identify sources of measurement uncertainty/error that introduce statistical uncertainty/error, that introduce systematic uncertainty/error, and that introduce both.
    2. Suggest approaches to reducing or constraining measurement uncertainty (both statistical and systematic).

Additional Content

You must be logged in to see this content.