Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

3.1 Probabilistic Reasoning

From Sense & Sensibility & Science
Revision as of 21:37, 1 August 2023 by Gpe (talk | contribs)

Using meta-judgments of the likelihood that your best judgment is right— how confident you are—enables decisions that take uncertainty into account.



The Lesson in Context

After introducing the concept of scientific uncertainty in previous lessons, we now teach the students that this uncertainty permeates all discussions of facts. Every factual claim should inherently carry a level of confidence as a percentage. It allows scientists to be open to the possibility that they may be wrong, while still being able to meaningfully discuss and compare the validity of factual statements. We aim to teach students that this way of thinking is important in daily life, often in the context of risk assessment, as well as in common discourse about social issues.

1.1 Introduction and When Is Science Relevant
  • Facts vs. values: Since credence levels can only be assigned to factual statements, it is important to first distinguish between statements of fact and statements of value.
2.2 Systematic and Statistical Uncertainty
  • Measurements in the real world are imperfect, and measurement uncertainties/errors can be studied and quantified. This translates to a confidence interval for every measurement result, i.e. "We are x% confident that the true value lies within this interval."
3.2 Calibration of Credence Levels
  • This lesson will follow up on the current one by teaching students how to calculate the calibration, or quality, of their credence, noticing and quantifying both underconfidence and overconfidence.
4.1 Signal and Noise
  • In that lesson, students explore how the signal they are looking for in data can be difficult to find amid the noise (random variation, random error, imperfect measurements, etc.). Because data is a mix of signal and noise, inferences from data tend to have some degree of uncertainty, which may be usefully quantified using credence levels or probabilities.
4.2 Finding Patterns in Random Noise
  • Since spurious patterns are expected to arise from random noise alone, any claim of actual pattern must carry with it a level of confidence that it is not due to random noise.
  • p-value: The probability that the observed pattern is due to random noise. In other words, one minus the p-value gives the level of confidence that the observed pattern is not due to random noise.
5.1 False Positives and Negatives
  • Since every binary test has a certain rate of false positives and false negatives, the result of such a test should only be understood as a recommendation of odds or risks, rather than a conclusive determination. Successive test results help one adjust their belief as well as their confidence level in that belief, e.g. whether one is suffering from a disease.


Takeaways

After this lesson, students should

  1. Recognize that every claim comes with some degree of uncertainty.
  2. Learn the function/utility of scientific expressions of uncertainty.
  3. Understand that because every proposition comes with a degree of uncertainty:
    1. Partial and probabilistic information still has value.
    2. Back-up plans are important since no information is absolutely certain.
    3. Evaluation of expertise and authority should be more directed towards accurately assigning confidence levels, rather than assuming a true expert would be "right" every single time.
    4. Scientific culture primarily uses a language of probabilities, and sometimes even well-confirmed facts turn out to be incomplete or not true in every single case.
    5. Even correctly done science will obtain incorrect results some of the time.

Credence

Level of confidence that a claim is true, from 0 (0 chance it is true) to 1 (100% certain it is true).

Confidence

Essentially a synonym for credence, as in "level of confidence," instead of colloquial meaning, "state of having a lot of confidence."

Accuracy

How frequently one is correct; proximity to a true value.

Calibration of Confidence

How closely confidence and accuracy correspond; that is, how accurate a person or system is at estimating the probability that they are correct.
  • p-value
The statistic used most often as a measure of statistical significance. The probability of getting a result as extreme or more if in fact the hypothesis is false, simply through random noise. The typical cut-off for a statistically significant p-value is p < .05.

Statistical Significance

How unlikely a given set of results would be if the null hypothesis were true (i.e. if the hypothesized effect did not actually exist).

A common misconception is that p-values are the probability of the hypothesis being false. This is not quite the same thing.


Where Credence Levels Come From

Credence levels can come from lots of places.
  • Past experience (how confident I feel about a statement of fact)
  • Instrumental uncertainties
  • Natural statistical variance (e.g. people come in different heights, wind speed is different across a town)

Where People Use Credence Levels

There are lots of familiar places where people already use or encounter credence levels.
  • Weather forecasts, specifically chances of rain.
  • Using polls to predict elections, phrased in terms of odds for betting (e.g. 50 to 1).
  • Making decisions based on a probabilistic risk assessment.
  • Credence levels predicting the probabilities of natural disasters within specific time frames (earthquakes, floods, wildfires etc.), and using these to make decisions about disaster preparation.
  • Use credence levels about getting into various colleges to decide what to use as a safety school.

My aunt still caught COVID even after taking the vaccine, therefore it's not true that the vaccine prevents COVID.

When phrased in terms of risks, the effectiveness of a vaccine is understood as the reduced probability that one would contract the disease after taking the vaccine, rather than total protection. Partial, probabilistic, improvements still have tangible causal consequences. This is expanded upon when we go over the distinction between singular and general causation.

Scientists used to say this particle is massless (its mass equals zero), but now they say it has a slightly non-zero mass. Their original measurement must've been wrong.

Despite how it's commonly described (sometimes even by scientists themselves), scientists don't typically measure the value of a quantity. Instead, they always measure a quantity to within some range of values, and then they say how confident they are that the true value falls within that range (marked with ± or error bars).

Useful Resources





Recommended Outline

Before Class

Print out the handouts.

During Class

5 Minutes Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc.
18 Minutes Have the students fill out the credence level warmup survey.
49 Minutes Conduct both probabilistic discussions.
8 Minutes Have the students make their future credence levels predictions.

Lesson Content

Credence Level Warm-up

The students will fill out a survey in class where they answer numerous questions of fact and assign a credence level to each of their answers. (They are not allowed to search the web.) This should get them comfortable with the process of assigning credence levels to claims and hopefully catch any overconfidence.

Instructions

1 Minute Hand out copies of the survey to each student.
5 Minutes Have students guess true/false for statements on a worksheet, and give a credence level for each answer.
3 Minutes Read the answers to the statements, and have students mark their own answers as true or false.
3 Minutes Instruct the students to take the average of their credence levels and the percent they got correct, and compare them.
6 Minutes Ask the followup questions after the survey.

Questions for the Class

  1. Did anyone rate something 100% confident and get it wrong?
  2. Raise your hand if you were overall underconfident (got more right than you expected).
  3. Raise your hand if you were overall overconfident (got more wrong than you expected).

You may allude to the idea of calibration of credence levels and how to adjust one's confidence in light of past experiences, which will be discussed in 3.2 Calibration of Credence Levels.

Survey Statements and Answers

In medieval Europe, the majority of people of your age and gender died before they reached the age of 50.

False. Expected lifespan was so low largely because of high infant and child mortality; if you made it to 17 or 18, you'd probably make it past 50.


The Northern Mockingbird, which mimics the songs of other birds, can learn as many as 200 distinct songs.

True. They can also learn to mimic car alarms, barking dogs, and ringtones.


Diamonds can be damaged by acid.

False. Diamonds are unaffected by acid and it is sometimes used to clean them.


The oldest tree is over 5,000 years old.

True. Bristlecone pines have been found that are over 5,000 years old.


Derby Street is north of Stuart Street in Berkeley.

True.


The heaviest recorded rabbit weighs over 45 pounds (20 kg).

True. The current record holder, a Flemish giant called Darius, weights 49 pounds (22 kg) and is 4 feet, 4 inches long.


The Republican party currently has a majority in the US Senate.

False. Although the Republicans have 50 senators, the Dems only 48, there are two independents who caucus with the Dems, and the tie is broken by VP Kamala Harris, a Democrat.


More Americans were killed last year by automobiles than by firearms.

True.


Over 50 million people died in WWI.

False. About 20 million people died in WWI.


If an eyewitness to a burglary confidently picks the suspect out of a lineup, then the perpetrator is always the person who the witness saw.

False. Eyewitnesses are notoriously unreliable, but their testimony is frequently overrated by jurors and judges.

Probabilistic Discussions

This activity lets students practice being mindful of how confident they are when they make various claims of fact, such as in a discussion on a social issue.

In small groups, the students will discuss two topics while assigning credence levels to each factual claim that they make. Students will learn to distinguish between claims of fact and claims of value and be aware of their level of confidence in claims of fact. Each group will have one designated notetaker and one designated moderator, who will not be actively discussing. After the first discussion, you'll have a brief discussion with the entire class. Then the students will do one more discussion with roles switched so that everyone gets a chance to do the discussing.

Instructions

2 Minutes Have the class select a topic to discuss. Then briefly explain the task, emphasizing the importance of the distinction between claims of fact and claims of value, and that credence levels (0-100%) are only assigned to claims of fact. Let each group decide on who takes on the following roles.
  1. Pedant: Reminds people to give a credence level for every factual claim. Also moderates to make sure everyone participates.
  2. Other members of the group: Discusses the topic, stating the credence level for each claim of fact they make.
15 Minutes Have the students discuss the first topic.
10 Minutes Run the interlude discussion.
1 Minute The class picks the second topic to discuss from the same list.
15 Minutes Have the students discuss the second topic.
6 Minutes Ask the final question.

Topic Selection List

Before explaining the task, the class should pick from the following list of topics to be discussed.

  1. Does standardized testing in the schools help or hurt education?
  2. Is social media a net harm or benefit for society?
  3. Is democracy the best form of government?
  4. Should violent video games be regulated?

First Topic

Have the students do the following.

  1. Choose a pedant for this activity.
  2. Discuss for 15 minutes.

Interlude

Ask the students the following questions.

  1. How did it feel to use credence levels? What did you notice?
  2. What was the typical range of credence levels you used? How often did people use 100%?
  3. Share some examples of factual claims with credence levels. How did you come up with these credence levels? What evidence would make you more sure? What evidence would make you less sure?
  4. Was it difficult to separate out factual claims from value claims? Why or why not?

Second Topic

Have the students do the following.

  1. Switch up roles so that the initial pedant participates in the discussion.
  2. Discuss the second chosen topic for 15 minutes.

Final Question

Did you notice anything different in the discussion of your second topic? Did it get easier to use credence levels with practice?

Future Predictions

The back of the handouts from the first activity should have a form for students to make predictions that will happen in the next two days (before the next class). They'll make seven predictions with different credence levels, which will be revisited in the next lesson to assess the calibration of the whole class. Have the students hold on to their handouts so that they'll be available for reference in the next lesson.

Students will probably lose the handout. Have them also take pictures of it with their phones.