2.2 Systematic and Statistical Uncertainty: Difference between revisions
More actions
// Edit via Wikitext Extension for VSCode |
// Edit via Wikitext Extension for VSCode |
||
| Line 263: | Line 263: | ||
|description=Reading regarding FiveThirtyEight's model for the 2016 US presidential election.}} | |description=Reading regarding FiveThirtyEight's model for the 2016 US presidential election.}} | ||
Spend four minutes on each question. | Spend four minutes on each question. | ||
<ol start=1><li>Why didn't the huge volume of polling data during the 2016 election eliminate uncertainty?</li></ol> | |||
= | |||
Why didn't the huge volume of polling data during the 2016 election eliminate uncertainty? | |||
{{BoxAnswer|Because the polls shared systematic errors, more often reaching kinds of people who were more likely to vote for Clinton and missing the kinds of people who were more likely to vote for Trump.}} | {{BoxAnswer|Because the polls shared systematic errors, more often reaching kinds of people who were more likely to vote for Clinton and missing the kinds of people who were more likely to vote for Trump.}} | ||
= | <ol start=2><li>Why is it valuable to have multiple polls?</li></ol> | ||
Why is it valuable to have multiple polls? | |||
{{BoxAnswer|Each poll is a bit different, using slightly different strategies and reaching different sets of people, so they can cancel out some of each other's systematic uncertainties.}} | {{BoxAnswer|Each poll is a bit different, using slightly different strategies and reaching different sets of people, so they can cancel out some of each other's systematic uncertainties.}} | ||
= | <ol start=3><li>Why do polls tend to replicate each other's mistakes?</li></ol> | ||
Why do polls tend to replicate each other's mistakes? | |||
{{BoxAnswer|Although no two polls are identical, they tend to use similar strategies for reaching likely voters and so may replicate systematic uncertainties across polls. For example, polls that depend on people picking up their phones from unknown callers will systematically miss people who don't pick up such calls, who may be systematically different. Even with absolute best practices, pollsters can only reach people who agree to answer questions, and willingness to answer a poll may be negatively correlated with antisocial tendencies or distrust in institutions, which may affect voting preferences.}} | {{BoxAnswer|Although no two polls are identical, they tend to use similar strategies for reaching likely voters and so may replicate systematic uncertainties across polls. For example, polls that depend on people picking up their phones from unknown callers will systematically miss people who don't pick up such calls, who may be systematically different. Even with absolute best practices, pollsters can only reach people who agree to answer questions, and willingness to answer a poll may be negatively correlated with antisocial tendencies or distrust in institutions, which may affect voting preferences.}} | ||
= | <ol start=4><li>Why was FiveThirtyEight's prediction better than the predictions of other pundits?</li></ol> | ||
Why was FiveThirtyEight's prediction better than the predictions of other pundits? | |||
{{BoxAnswer|FiveThirtyEight took into account the historical accuracy of the individual polls. By noting systematic discrepancies from that historical record, they were able to use more appropriate error bars taking the uncertainties into account.}} | {{BoxAnswer|FiveThirtyEight took into account the historical accuracy of the individual polls. By noting systematic discrepancies from that historical record, they were able to use more appropriate error bars taking the uncertainties into account.}} | ||
= | <ol start=5><li>Why are state polls less reliable than national polls?</li></ol> | ||
Why are state polls less reliable than national polls? | |||
{{BoxAnswer|Because they usually have smaller sample sizes and there are fewer polls to balance out systematic errors from any given poll.}}{{NavCard|prev=2.1 Senses and Instrumentation|next=3.1 Probabilistic Reasoning}} | {{BoxAnswer|Because they usually have smaller sample sizes and there are fewer polls to balance out systematic errors from any given poll.}}{{NavCard|prev=2.1 Senses and Instrumentation|next=3.1 Probabilistic Reasoning}} | ||
[[Category:Lesson plans]] | [[Category:Lesson plans]] | ||
Revision as of 19:54, 15 August 2023

This topic explores the sources of error and uncertainty in data.
The Lesson in Context
This lesson teaches students that the inevitable imperfections of instrumental measurements of the real world can be quantified and studied in their own rights. They are categorized into statistical uncertainty and systematic uncertainty. We will teach them to be aware of the sources of uncertainty in each measurement and some elementary ways to mitigate them. This is illustrated by the human histogram activity, in which students can see statistical distributions and physically experience the effects of systematic uncertainty. We will also discuss how systematic uncertainties affected results of political polling in the 2016 US presidential election.
Takeaways
After this lesson, students should
- Realize that our contact with reality is often mediated by measurement and quantification. We need to be aware that every measurement comes with some degree of uncertainty (deviation from the "true" value in reality).
- Identify sources of measurement uncertainty/error that introduce statistical uncertainty/error, that introduce systematic uncertainty/error, and that introduce both.
- Understand how to use repeated measures to reduce statistical uncertainty.
- Recognize the difficulty of removing systematic uncertainty, and that the process of science involves creativity in identifying sources of systematic uncertainty and inventing strategies to reduce or eliminate them.
Students will likely keep asking "but how can I tell between statistical and systematic uncertainties", and the answer would be to offer as many diverse examples as possible. Also if you can reduce the uncertainty simply by collecting more data, it's statistical.
Statistical Uncertainty
Systematic Uncertainty
Accuracy
Precision
A measurement can be very precise but wrong/inaccurate (low statistical uncertainty but high systematic uncertainty), or it could have a large variance between subsequent measurements but average accurately to the true value (low systematic uncertainty but high statistical uncertainty).
Proxy
Because a proxy is not a direct measure of the quantity of interest, it is a place that systematic bias can creep in. For example, using self-report ratings of happiness as a measure of happiness may be subject to cultural differences and/or comparison effects.
Triangulation
Proxy Example
- When measuring crime rate, it is only possible to measure the rate of reported crime, making it a proxy of the true crime rate. Even when measuring temperature, it is only possible to observe the reading on an instrument, also a form of proxy.
Simple Systematic Uncertainty Example
- "It won't do us any good to average lots of test subjects' heights together if our tape measure got shrunk in the wash!"
Simple Statistical Uncertainty Example
- "Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. The people they are talking to might not be representative of the whole population. For example, older people can be more likely to pick up the phone and talk to pollsters, so there might be a systematic bias in that direction."
Polling Models
- "Indeed, [the 2016] election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for [systematic uncertainties]. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out." (Source)
Stiffness of Springs
- "The stiffness of many springs depends on their temperature. If you measure the stiffness of a spring many times, by compressing and decompressing it, the internal friction inside the spring may cause it to warm. You may see this by a systematic trend in your data set; for example, each data point in a data set will be smaller than the previous one." (Source)
Systematic Error in Life Sciences
- "Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error — cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture." This is a more complex life science example of systematic error.
Reducing Systematic Error
- "If we estimate the effect of a drug on weight by randomly assigning people to take the drug vs. not take it and then measure their weight after a year, we could subtract the average weight loss of drug-takers vs. non-drug-takers to get the effect size of the drug on weight loss. But the people know if they're taking a drug for weight loss, so there could be a placebo effect creating a systematic bias. So the better way to do the experiment is to give the control group sugar pills. Then we can be more confident that any weight loss is due to the drug, and not a systematic bias created by the placebo effect."
Sleep Questionnaire
- If you asked just a handful of random people on the street how much they slept the night before, the average of their answers could be quite different from the true average of the whole population due to random differences between people (statistical uncertainty). This can be improved by asking more people (say, hundreds or thousands). However, if you asked hundreds of random people on a college campus the same question, all of their answers could be skewed in one direction due to collective sleep deprivation (systematic uncertainty), which would not be improved by asking more college students.
Brightness of a Star
- Suppose you are an astronomer measuring the brightness of a star. The star twinkles due to random atmospheric fluctuations (statistical uncertainty), but the presence of the atmosphere itself, together with clouds, always reduces the brightness of the star (systematic uncertainty).
The words "uncertainty" and "error" mean that our instruments or measurement methods are somehow broken, deficient, or not to be trusted.
A single measurement can only have either systematic or statistical uncertainty.
Useful Resources
Recommended Outline
Before Class
- Prepare sheets of paper with ranges of heights for the Human Histogram activity. You may also want to scout a location and get some tape to hang the sheets with.
- Remind students to read the FiveThirtyEight article.
During Class
| 5 Minutes | Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc. |
| 2 Minutes | Have the students attempt the warm-up question. |
| 7 Minutes | Go through the concept review questions. |
| 20 Minutes | Walk the students through the FiveThirtyEight reading discussion. |
| 33 Minutes | Conduct the Human Histogram activity. |
| Remaining Time | This leaves 13 minutes of wiggle room for any activities that go long. If you have time at the end you can collect and answer any people's lingering questions. |
Lesson Content
Warm-up Question
Making lots of measurements and averaging them is a good way to reduce...
- Statistical Uncertainty
- Systematic Uncertainty
Concept Review Questions
Concept Review Question 1
Suppose you've discovered a new substance. You've taken it into the lab and you want to see what temperature it melts at. But, you're worried that the thermometer is miscalibrated. Would this be a source of statistical or systematic uncertainty in measurements that use this thermometer? How might you figure this out and reduce this uncertainty?
Miscalibration of instrument is a source of systematic uncertainty. You can test it by putting the thermometer into just boiling water and see if it's 100 degrees Celsius, and into just freezing water and see if it's 0 degrees Celsius. OR you can compare it to another thermometer.
Concept Review Question 2
Suppose you're running a study of the effects of class size on student learning. You find that students at a school with average class size of 40 perform worse than students at a neighboring school with average class size of 30. Both schools have 3000 students. Should you be more worried about systematic or statistical uncertainty? If you had abundant control and resources, what could you do to reduce each source of uncertainty that you can think of?
You should be more worried about systematic uncertainty because there are likely to be other differences between the two schools and the students. You can reduce systematic uncertainty by randomly assigning students to school (with bussing) and/or by randomly assigning classes to have 30 vs. 40 students, with equal numbers of small/large classrooms at each school. This takes out the confound of school (and the many variables that may differ between schools) with class size. You could reduce statistical uncertainty by doing the same thing at many schools, not only two.
Human Histogram
This activity aims to illustrate the effects of statistical and systematic uncertainties in data by having students participate as individual data points, forming a "human histogram" in which each student is actually a living data point.
Preparation
Finding a Location
If possible, find a location outdoors near your classroom. This is where the activity will be conducted. So, make sure that it's a place you could quickly migrate the students to during class. The activity can be done in the classroom if you don't find a good spot. But, space may be more cramped.
Make sheets of paper with large print or writing to denote the bins in the histogram. If you have 30 or so students then the following four categories should be sufficient. If you have 60-80 students then it may be best to have six or so bins instead of four. Make sure you have any scotch tape to hold up the signs indoors, rocks to hold them down outside, or anything else you may need in order to have the signs marked clearly.
| Imperial | Metric |
|---|---|
| < 5'2" | < 158 cm |
| 5'2" - 5'7" | 158 cm - 170 cm |
| 5'7" - 6' | 170 cm - 183 cm |
| > 6' | > 183 cm |
How the activity will be carried out heavily depends on the exact venue. Make sure to familiarize yourself with the venue and rehearse it ahead of time. During or just before class, place the papers corresponding to the "bins" on opposite sides of whatever space you're conducting the activity. There should be one set of bins for football fans and one for soccer fans. Make sure there's enough space that several students can stand at each piece of paper without having the bins bleed into each other.
Instructions
| 4 Minutes | If you're not doing the activity in your classroom, take this time to move all the students outside. |
| 13 Minutes | Conduct the "During the Game" part of the activity. |
| 16 Minutes | Conduct the "At the Party" part of the activity. |
During the Game
| 2 Minutes | Split the class into two halves. Let students choose whether they'd like to be football fans or soccer fans. |
| 3 Minutes | Tell students that they are attending a football or soccer game as per their preference. You want to know how many people of what height are at each game. Instruct the students to walk into the bin that contains their height (e.g. for someone 5'7", go into the 5'6" - 5'9" bin) on their own side. |
| 1 Minute | When it settles, write down which histogram skews taller. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set. |
| 3 Minutes | The instructor should now role play as the "alien researcher" who is studying human heights. Eyeball the height distributions of the two groups to see which group is taller on average than the other. The alien researcher says, "it looks like the [football/soccer] fans here are taller than the [soccer/football] fans, so [football/soccer] fans must in general be taller than [soccer/football] fans!" |
| 4 Minutes | Ask the students the final if this alien researcher's conclusion is justified? Why or why not? |
Answer
The alien is not justified in their conclusion, because they made a generalization about an overall population based on only very few samples. The differences in heights that we see in this small class are subject to statistical uncertainty.
Some students might not have a preference between football and soccer. Try to make the two halves more or less similar in size.
At the Party
| 3 Minutes | A second alien researcher decides to do the same study, but at the after parties. Now tell the story that football fans are going to a beach party, where they may wear flip flops or no shoes at all. On the other hand, soccer fans are going to a fancy hat ball, where they may wear fancy hats and/or high heels. |
| 3 Minutes | Instruct students to now imagine how tall they may be including their new apparel (extremely flamboyant for soccer fans) and walk into the bins containing their new height. Make it clear that soccer fans should join the bin indicating their height including the imaginary heels and hats, because the alien assumes the shoes and hats are part of the human organism. |
| 1 Minute | When it settles, write down which histogram skews taller and screenshot them/photograph. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set. |
| 3 Minutes | he second alien researcher says, "it looks like the soccer fans are taller than the football fans, so soccer fans must in general be taller than football fans!" |
| 6 Minutes | Ask the students the final if this alien researcher's conclusion is justified? Why or why not? |
Answer
The alien is not justified in their conclusion, not only because of statistical uncertainty, but also because there is now a major source of systematic uncertainty—the hats and shoes on the soccer fans.
Conclude that statistical uncertainty could go both ways simply due to the randomness in the selection of samples (such as heights of students in this section), and it can be improved by increasing the sample size (e.g. the whole SSS class).
Systematic uncertainties cause the data to be skewed in one direction, and cannot be removed just by increasing sample size. To remove this effect, the alien would have to understand human physiology—that hats and shoes are not part of the human body—and account for that.
We did this activity with only the small number of students in a single discussion section. Did it work? How would it have worked better if we did it with the entire class? You can use this as a teaching opportunity to show how systematic and statistical errors manifest in small sample sizes.
FiveThirtyEight Reading Discussion
Why FiveThirtyEight Gave Trump a Better Chance Than Almost Anyone Else
Reading regarding FiveThirtyEight's model for the 2016 US presidential election.
Spend four minutes on each question.
- Why didn't the huge volume of polling data during the 2016 election eliminate uncertainty?
Because the polls shared systematic errors, more often reaching kinds of people who were more likely to vote for Clinton and missing the kinds of people who were more likely to vote for Trump.
- Why is it valuable to have multiple polls?
Each poll is a bit different, using slightly different strategies and reaching different sets of people, so they can cancel out some of each other's systematic uncertainties.
- Why do polls tend to replicate each other's mistakes?
Although no two polls are identical, they tend to use similar strategies for reaching likely voters and so may replicate systematic uncertainties across polls. For example, polls that depend on people picking up their phones from unknown callers will systematically miss people who don't pick up such calls, who may be systematically different. Even with absolute best practices, pollsters can only reach people who agree to answer questions, and willingness to answer a poll may be negatively correlated with antisocial tendencies or distrust in institutions, which may affect voting preferences.
- Why was FiveThirtyEight's prediction better than the predictions of other pundits?
FiveThirtyEight took into account the historical accuracy of the individual polls. By noting systematic discrepancies from that historical record, they were able to use more appropriate error bars taking the uncertainties into account.
- Why are state polls less reliable than national polls?
Because they usually have smaller sample sizes and there are fewer polls to balance out systematic errors from any given poll.