Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

3.2 Calibration of Credence Levels: Difference between revisions

From Sense & Sensibility & Science
Created page with "{{subst:Lesson plan}}"
 
Line 14: Line 14:
[Link to instructional video]
[Link to instructional video]


More details: [link to topic on SSS website]
[https://sensesensibilityscience.berkeley.edu/topic/13 More details]


After this lesson, students should
After this lesson, students should
# [What they should understand, feel, learn, etc.]
# Be wary of high levels of confidence.
# Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars, 3σ).
# Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration.


=== Definitions ===
=== Definitions ===


* '''Term'''
* '''Confidence interval'''
*: Explanation.
*: A range of values within which the true value of interest lies with a specified probability (typically 95%, meaning there is a 5% chance that the true value falls outside of this range).
* '''Error bars'''
*: Smaller bars on a graph that show the range of likely true values around the observed value, typically a 95% confidence interval, or the observed value ± the standard error or standard deviation (equivalent to 68% confidence interval).


=== Examples ===
=== Examples ===


* [item 1]
* In the figure below, lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. ([https://www.researchgate.net/publication/228212908_Insightful_or_Wishful_Lawyers'_Ability_to_Predict_Case_Outcomes Source]) [[File:Credence Level Lawyers.png|center|400px]]
* In this figure, calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. ([https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1365-2648.2010.05437.x Source]) {{Caution|Note that the probabilities are all greater than 50%. This is fine since low likelihood can be represented on a plot like this whenever there’s a binary prediction about whether something will happen. You just “flip” the prediction from whether something ''will'' happen to whether something will not happen.}} [[File:Credence Level Nurses.png|center|400px]]
* It is possible (and expected) for singular election forecasts to turn out to be wrong some of the time. It is fine as long as the predictions within that credence range are overall well calibrated. It is also impossible to calibrate the credence of a singular prediction, it is only possible for an aggregate of predictions. ([https://fivethirtyeight.com/features/the-polls-werent-great-but-thats-pretty-normal/ Source])


=== Common Misconceptions ===
=== Common Misconceptions ===


* ''Mistaken quote''
* ''He seems super confident, and she said she was only 85% sure, so we should trust him over her.''
*: Explanation.
*: It is more important to have the self-awareness of how often they are wrong, than to always insist they are right. If possible, use the outcomes of their past predictions to evaluate the calibration of their confidence before placing trust in them.
* ''Am I well calibrated in this particular prediction?''
*: The question doesn’t really make sense. You can’t have a calibration for a single prediction. A calibration level is only meaningfully defined over a collection of predictions for which you can see how many came ultimately true.


== Context ==
== Context ==

Revision as of 03:33, 27 November 2021


Fill out the learning goals and definitions.
Add links in learning goals.
Write the "context" for how this lesson connects with others.
Edit previous lesson to see if there's any setup for this lesson that needs to be in the "During Class" or "After Class" section of the last.
Fill out the lesson content.
Fill out or delete anything with "[...]" remaining.

Learning Goals

[Link to PlayPosit]

[Link to instructional video]

More details

After this lesson, students should

  1. Be wary of high levels of confidence.
  2. Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars, 3σ).
  3. Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration.

Definitions

  • Confidence interval
    A range of values within which the true value of interest lies with a specified probability (typically 95%, meaning there is a 5% chance that the true value falls outside of this range).
  • Error bars
    Smaller bars on a graph that show the range of likely true values around the observed value, typically a 95% confidence interval, or the observed value ± the standard error or standard deviation (equivalent to 68% confidence interval).

Examples

  • In the figure below, lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. (Source)
  • In this figure, calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. (Source)
    Note that the probabilities are all greater than 50%. This is fine since low likelihood can be represented on a plot like this whenever there’s a binary prediction about whether something will happen. You just “flip” the prediction from whether something will happen to whether something will not happen.
  • It is possible (and expected) for singular election forecasts to turn out to be wrong some of the time. It is fine as long as the predictions within that credence range are overall well calibrated. It is also impossible to calibrate the credence of a singular prediction, it is only possible for an aggregate of predictions. (Source)

Common Misconceptions

  • He seems super confident, and she said she was only 85% sure, so we should trust him over her.
    It is more important to have the self-awareness of how often they are wrong, than to always insist they are right. If possible, use the outcomes of their past predictions to evaluate the calibration of their confidence before placing trust in them.
  • Am I well calibrated in this particular prediction?
    The question doesn’t really make sense. You can’t have a calibration for a single prediction. A calibration level is only meaningfully defined over a collection of predictions for which you can see how many came ultimately true.

Context

[Explanation of how the current topic into the larger context of the course by explaining how the previous topic(s) relate to the current topic and leaving cliffhangers for future topics where relevant throughout the lesson]

Before

X.X Lesson Name
Explanation.
  • Itemized explanation.

After

X.X Lesson Name
Explanation.
  • Itemized explanation.

Recommended Outline

Before Class

  • [Any essential logistical things that need to be done for this class]
  • Prepare a seating chart.
  • Review PlayPosit and discussion questions and ask faculty, Gabriel, or Emlen any questions you have.
  • (Optional) Prepare a presentation.

During Class

  • (5 min) Come up with some fun way to assign the roles of spokesperson and notetaker (e.g. earliest birthday in the year, lives furthest from campus). Remind them of the responsibilities of these roles.
  • ([#] min) [Description of some module from #Lesson Content]
  • (5 min) Collect questions for plenary.

After Class

  • [Any essential logistical things that need to be done as followup for this class]
  • Collect answers from notetakers for the forum / plenary.

Lesson Content

Clicker Question

Question

Correct answer.
  1. First option
  2. Second option
  3. Third option
  4. Fourth option

Activity 1: Name

[Brief description of and motivation for the activity]

Common misconceptions and any useful tricks, tips, guidelines, or other background

Instructions

Discussion Questions

  1. Question 1
    1. Subquestion a
      Possible misconception that may need to be corrected and clarified.
      Intended answer to the above question.
    2. Subquestion b

Collect Questions for Plenary

(5 min) Collect remaining questions from the students for faculty in plenary (can be questions for clarification, extension, discussion, etc.), and add [ here].