Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

2.2 Systematic and Statistical Uncertainty: Difference between revisions

From Sense & Sensibility & Science
No edit summary
// Edit via Wikitext Extension for VSCode
Tag: Reverted
Line 25: Line 25:


== Takeaways ==
== Takeaways ==
 
QQ
<tabber>
<tabber>



Revision as of 22:10, 21 August 2023

This topic explores the sources of error and uncertainty in data.



The Lesson in Context

This lesson teaches students that the inevitable imperfections of instrumental measurements of the real world can be quantified and studied in their own rights. They are categorized into statistical uncertainty and systematic uncertainty. We will teach them to be aware of the sources of uncertainty in each measurement and some elementary ways to mitigate them. This is illustrated by the human histogram activity, in which students can see statistical distributions and physically experience the effects of systematic uncertainty. We will also discuss how systematic uncertainties affected results of political polling in the 2016 US presidential election.

1.2 Shared Reality and Modeling
  • It is inevitable that our experience or measurement of the external reality is imperfect. This lesson's concepts help to quantify these imperfections.
2.1 Senses and Instrumentation
  • No instrument is perfect. Systematic and statistical uncertainties help quantify these imperfections and allow us to compare two different instruments or methods of measurement.
6.1 Correlation and Causation
  • Randomized assignment is one way to remove the systematic uncertainty by making sure that the intervention and control groups are not correlated with some other variable related to the method of assignment itself, e.g. a male vs. female group in a drug trial.
  • Placebo effect is a systematic uncertainty in the measurement of the effectiveness of a treatment. Therefore, we must "subtract" the effect of the placebo treatment from the effect of the real treatment.


Takeaways

QQ

After this lesson, students should

  1. Realize that our contact with reality is often mediated by measurement and quantification. We need to be aware that every measurement comes with some degree of uncertainty (deviation from the "true" value in reality).
  2. Identify sources of measurement uncertainty/error that introduce statistical uncertainty/error, that introduce systematic uncertainty/error, and that introduce both.
  3. Understand how to use repeated measures to reduce statistical uncertainty.
  4. Recognize the difficulty of removing systematic uncertainty, and that the process of science involves creativity in identifying sources of systematic uncertainty and inventing strategies to reduce or eliminate them.

Students will likely keep asking "but how can I tell between statistical and systematic uncertainties", and the answer would be to offer as many diverse examples as possible. Also if you can reduce the uncertainty simply by collecting more data, it's statistical.


Statistical Uncertainty

Differences between reality and measurements on the basis of random imprecisions or "noise."

Systematic Uncertainty

Differences between reality and measurements that skew results in one direction.

Accuracy

How close the measured value is to the true value.

Precision

How similar are all the measured values of the same thing (consistency).

A measurement can be very precise but wrong/inaccurate (low statistical uncertainty but high systematic uncertainty), or it could have a large variance between subsequent measurements but average accurately to the true value (low systematic uncertainty but high statistical uncertainty).

Proxy

An observable measurement used to approximate a quantity that is not directly observable or measurable. E.g. people's ratings of agreement with the statement "I am happy" on a scale from 1 to 7, is a measure (proxy) of their happiness, or using zip code as a proxy for socio-economic status.

Because a proxy is not a direct measure of the quantity of interest, it is a place that systematic bias can creep in. For example, using self-report ratings of happiness as a measure of happiness may be subject to cultural differences and/or comparison effects.

Triangulation

A method of dealing with systematic errors inherent in a type of measurement by attempting to measure the same phenomenon in multiple different ways or through lots of different proxies.


Proxy Example

When measuring crime rate, it is only possible to measure the rate of reported crime, making it a proxy of the true crime rate. Even when measuring temperature, it is only possible to observe the reading on an instrument, also a form of proxy.

Simple Systematic Uncertainty Example

"It won't do us any good to average lots of test subjects' heights together if our tape measure got shrunk in the wash!"

Simple Statistical Uncertainty Example

"Sure, those polls all claim to be accurate within three percentage points, but they just mean that their statistical accuracy is that good. The people they are talking to might not be representative of the whole population. For example, older people can be more likely to pick up the phone and talk to pollsters, so there might be a systematic bias in that direction."

Polling Models

"Indeed, [the 2016] election has demonstrated, quite emphatically, that none of the polling models out there have adequately controlled for [systematic uncertainties]. Unless you understand and quantify your systematic errors—and you can't do that if you don't understand how your polling might be biased—election forecasts will suffer from the GIGO problem: garbage in, garbage out." (Source)

Stiffness of Springs

"The stiffness of many springs depends on their temperature. If you measure the stiffness of a spring many times, by compressing and decompressing it, the internal friction inside the spring may cause it to warm. You may see this by a systematic trend in your data set; for example, each data point in a data set will be smaller than the previous one." (Source)

Systematic Error in Life Sciences

"Biologists often test cancer drugs on cell lines. Cell lines are cell cultures (groups of living cells grown under controlled conditions, generally outside their natural environment) with a uniform genetic makeup. Conclusions about all cells of the cell type made from measurements or experiments performed on cell lines suffer from systematic error — cells in the body do not have a completely uniform genetic makeup and exist in conditions vastly different from a cell culture." This is a more complex life science example of systematic error.

Reducing Systematic Error

"If we estimate the effect of a drug on weight by randomly assigning people to take the drug vs. not take it and then measure their weight after a year, we could subtract the average weight loss of drug-takers vs. non-drug-takers to get the effect size of the drug on weight loss. But the people know if they're taking a drug for weight loss, so there could be a placebo effect creating a systematic bias. So the better way to do the experiment is to give the control group sugar pills. Then we can be more confident that any weight loss is due to the drug, and not a systematic bias created by the placebo effect."

Sleep Questionnaire

If you asked just a handful of random people on the street how much they slept the night before, the average of their answers could be quite different from the true average of the whole population due to random differences between people (statistical uncertainty). This can be improved by asking more people (say, hundreds or thousands). However, if you asked hundreds of random people on a college campus the same question, all of their answers could be skewed in one direction due to collective sleep deprivation (systematic uncertainty), which would not be improved by asking more college students.

Brightness of a Star

Suppose you are an astronomer measuring the brightness of a star. The star twinkles due to random atmospheric fluctuations (statistical uncertainty), but the presence of the atmosphere itself, together with clouds, always reduces the brightness of the star (systematic uncertainty).

The words "uncertainty" and "error" mean that our instruments or measurement methods are somehow broken, deficient, or not to be trusted.

These words describe the inevitable and perfectly acceptable gap between measurement and reality.

A single measurement can only have either systematic or statistical uncertainty.

Every measurement can come with systematic and statistical uncertainty, often with multiple sources to different degrees.

<restricted>

Useful Resources






Recommended Outline

Before Class

During Class

5 Minutes Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc.
2 Minutes Have the students attempt the warm-up question.
7 Minutes Go through the concept review questions.
20 Minutes Walk the students through the FiveThirtyEight reading discussion.
33 Minutes Conduct the Human Histogram activity.
Remaining Time This leaves 13 minutes of wiggle room for any activities that go long. If you have time at the end you can collect and answer any people's lingering questions.

Lesson Content

Warm-up Question

Making lots of measurements and averaging them is a good way to reduce...

  1. Statistical Uncertainty
  2. Systematic Uncertainty

Concept Review Questions

  1. Suppose you've discovered a new substance. You've taken it into the lab and you want to see what temperature it melts at. But, you're worried that the thermometer is miscalibrated. Would this be a source of statistical or systematic uncertainty in measurements that use this thermometer? How might you figure this out and reduce this uncertainty?

Miscalibration of instrument is a source of systematic uncertainty. You can test it by putting the thermometer into just boiling water and see if it's 100 degrees Celsius, and into just freezing water and see if it's 0 degrees Celsius. OR you can compare it to another thermometer.

  1. Suppose you're running a study of the effects of class size on student learning. You find that students at a school with average class size of 40 perform worse than students at a neighboring school with average class size of 30. Both schools have 3000 students. Should you be more worried about systematic or statistical uncertainty? If you had abundant control and resources, what could you do to reduce each source of uncertainty that you can think of?

You should be more worried about systematic uncertainty because there are likely to be other differences between the two schools and the students. You can reduce systematic uncertainty by randomly assigning students to school (with bussing) and/or by randomly assigning classes to have 30 vs. 40 students, with equal numbers of small/large classrooms at each school. This takes out the confound of school (and the many variables that may differ between schools) with class size. You could reduce statistical uncertainty by doing the same thing at many schools, not only two.

Human Histogram

This activity aims to illustrate the effects of statistical and systematic uncertainties in data by having students participate as individual data points, forming a "human histogram" in which each student is actually a living data point.

Preparation

Finding a Location

If possible, find a location outdoors near your classroom. This is where the activity will be conducted. So, make sure that it's a place you could quickly migrate the students to during class. The activity can be done in the classroom if you don't find a good spot. But, space may be more cramped.

Make sheets of paper with large print or writing to denote the bins in the histogram. If you have 30 or so students then the following four categories should be sufficient. If you have 60-80 students then it may be best to have six or so bins instead of four. Make sure you have any scotch tape to hold up the signs indoors, rocks to hold them down outside, or anything else you may need in order to have the signs marked clearly.

Imperial Metric
< 5'2" < 158 cm
5'2" - 5'7" 158 cm - 170 cm
5'7" - 6' 170 cm - 183 cm
> 6' > 183 cm

How the activity will be carried out heavily depends on the exact venue. Make sure to familiarize yourself with the venue and rehearse it ahead of time. During or just before class, place the papers corresponding to the "bins" on opposite sides of whatever space you're conducting the activity. There should be one set of bins for football fans and one for soccer fans. Make sure there's enough space that several students can stand at each piece of paper without having the bins bleed into each other.

Instructions

4 Minutes If you're not doing the activity in your classroom, take this time to move all the students outside.
13 Minutes Conduct the "During the Game" part of the activity.
16 Minutes Conduct the "At the Party" part of the activity.
During the Game
2 Minutes Split the class into two halves. Let students choose whether they'd like to be football fans or soccer fans.
3 Minutes Tell students that they are attending a football or soccer game as per their preference. You want to know how many people of what height are at each game. Instruct the students to walk into the bin that contains their height (e.g. for someone 5'7", go into the 5'6" - 5'9" bin) on their own side.
1 Minute When it settles, write down which histogram skews taller. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set.
3 Minutes The instructor should now role play as the "alien researcher" who is studying human heights. Eyeball the height distributions of the two groups to see which group is taller on average than the other. The alien researcher says, "it looks like the [football/soccer] fans here are taller than the [soccer/football] fans, so [football/soccer] fans must in general be taller than [soccer/football] fans!"
4 Minutes Ask the students the final if this alien researcher's conclusion is justified? Why or why not?

Answer

The alien is not justified in their conclusion, because they made a generalization about an overall population based on only very few samples. The differences in heights that we see in this small class are subject to statistical uncertainty.

Some students might not have a preference between football and soccer. Try to make the two halves more or less similar in size.

At the Party
3 Minutes A second alien researcher decides to do the same study, but at the after parties. Now tell the story that football fans are going to a beach party, where they may wear flip flops or no shoes at all. On the other hand, soccer fans are going to a fancy hat ball, where they may wear fancy hats and/or high heels.
3 Minutes Instruct students to now imagine how tall they may be including their new apparel (extremely flamboyant for soccer fans) and walk into the bins containing their new height. Make it clear that soccer fans should join the bin indicating their height including the imaginary heels and hats, because the alien assumes the shoes and hats are part of the human organism.
1 Minute When it settles, write down which histogram skews taller and screenshot them/photograph. There should not be a big difference between the two halves. If your class has multiple discussion sections, hold onto this data. The results from every section will should be put together in one large data set.
3 Minutes he second alien researcher says, "it looks like the soccer fans are taller than the football fans, so soccer fans must in general be taller than football fans!"
6 Minutes Ask the students the final if this alien researcher's conclusion is justified? Why or why not?

Answer

The alien is not justified in their conclusion, not only because of statistical uncertainty, but also because there is now a major source of systematic uncertainty—the hats and shoes on the soccer fans.

Conclude that statistical uncertainty could go both ways simply due to the randomness in the selection of samples (such as heights of students in this section), and it can be improved by increasing the sample size (e.g. the whole SSS class).

Systematic uncertainties cause the data to be skewed in one direction, and cannot be removed just by increasing sample size. To remove this effect, the alien would have to understand human physiology—that hats and shoes are not part of the human body—and account for that.

We did this activity with only the small number of students in a single discussion section. Did it work? How would it have worked better if we did it with the entire class? You can use this as a teaching opportunity to show how systematic and statistical errors manifest in small sample sizes.

FiveThirtyEight Reading Discussion

Spend four minutes on each question.

  1. Why didn't the huge volume of polling data during the 2016 election eliminate uncertainty?

Because the polls shared systematic errors, more often reaching kinds of people who were more likely to vote for Clinton and missing the kinds of people who were more likely to vote for Trump.

  1. Why is it valuable to have multiple polls?

Each poll is a bit different, using slightly different strategies and reaching different sets of people, so they can cancel out some of each other's systematic uncertainties.

  1. Why do polls tend to replicate each other's mistakes?

Although no two polls are identical, they tend to use similar strategies for reaching likely voters and so may replicate systematic uncertainties across polls. For example, polls that depend on people picking up their phones from unknown callers will systematically miss people who don't pick up such calls, who may be systematically different. Even with absolute best practices, pollsters can only reach people who agree to answer questions, and willingness to answer a poll may be negatively correlated with antisocial tendencies or distrust in institutions, which may affect voting preferences.

  1. Why was FiveThirtyEight's prediction better than the predictions of other pundits?

FiveThirtyEight took into account the historical accuracy of the individual polls. By noting systematic discrepancies from that historical record, they were able to use more appropriate error bars taking the uncertainties into account.

  1. Why are state polls less reliable than national polls?

Because they usually have smaller sample sizes and there are fewer polls to balance out systematic errors from any given poll.

</restricted>