3.2 Calibration of Credence Levels: Difference between revisions
More actions
// via Wikitext Extension for VSCode |
// via Wikitext Extension for VSCode |
||
| (One intermediate revision by the same user not shown) | |||
| Line 1: | Line 1: | ||
{{Cover|3.2 Calibration of Credence Levels}} | {{Cover|3.2 Calibration of Credence Levels}} | ||
How do we know what the appropriate degree of confidence for any statement of fact should be? Even experts often suffer from overconfidence. We introduce practical techniques to calibrate appropriate levels of confidence as well as psychological attitudes that motivate one to improve one's calibration. | How do we know what the appropriate degree of confidence for any statement of fact should be? Even experts often suffer from overconfidence. We introduce practical techniques to calibrate appropriate levels of confidence as well as psychological attitudes that motivate one to improve one's calibration. | ||
== The Lesson in Context == | == The Lesson in Context == | ||
<!-- Always begin section with a description of this lesson in relation to the course as a whole. --> | <!-- Always begin section with a description of this lesson in relation to the course as a whole. --> | ||
This lesson continues the previous lesson, [[3.1 Probabilistic Reasoning]], and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the [[3.1 Probabilistic Reasoning#Future Predictions|Future Predictions]] activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration. | This lesson continues the previous lesson, [[3.1 Probabilistic Reasoning]], and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the [[3.1 Probabilistic Reasoning#Future Predictions|Future Predictions]] activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration. | ||
<!-- Expandable section relating this lesson to other lessons. --> | <!-- Expandable section relating this lesson to other lessons. --> | ||
{{Expand|Relation to Other Lessons| | {{Expand|Relation to Other Lessons| | ||
| Line 29: | Line 25: | ||
}} | }} | ||
== Takeaways == | == Takeaways == | ||
<tabber> | <tabber> | ||
|-|Learning Goals= | |-|Learning Goals= | ||
After this lesson, students should | After this lesson, students should | ||
<!-- Learning goals are written as a numbered list. --> | <!-- Learning goals are written as a numbered list. --> | ||
| Line 39: | Line 32: | ||
# Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars). | # Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars). | ||
# Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration. | # Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration. | ||
|-|Definitions= | |-|Definitions= | ||
<!-- Definitions must be written with the Definition and Subdefinition templates. --> | <!-- Definitions must be written with the Definition and Subdefinition templates. --> | ||
{{Definition | {{Definition | ||
| Line 47: | Line 38: | ||
|A range of values within which we believe the true value lies, with some credence level. | |A range of values within which we believe the true value lies, with some credence level. | ||
|comments= | |comments= | ||
{{BoxTip | |||
|Confidence interval, when describing an instrumental measurement, corresponds to the statistical uncertainty. | |Confidence interval, when describing an instrumental measurement, corresponds to the statistical uncertainty. | ||
}} | |||
{{BoxCaution | |||
|When one presents a scientific measurement, they are usually not presenting a ''specific'' value. They're presenting a range that the value is within and the likelihood that the true value is within that range. | |When one presents a scientific measurement, they are usually not presenting a ''specific'' value. They're presenting a range that the value is within and the likelihood that the true value is within that range. | ||
}} | |||
{{BoxCaution | |||
|Different fields have different standards for confidence intervals. Physicists typically choose confidence intervals within which they are 68% sure the true value lies. Psychologists typically use 95%.}} | |Different fields have different standards for confidence intervals. Physicists typically choose confidence intervals within which they are 68% sure the true value lies. Psychologists typically use 95%.}} | ||
{{BoxCaution | |||
|There are other related terms like "standard deviation," "<math>\sigma</math>," and "standard error." Try to avoid this jargon. If students ask about these terms then discuss it outside of class and be mindful of students that haven't had a statistics class.}} | |There are other related terms like "standard deviation," "<math>\sigma</math>," and "standard error." Try to avoid this jargon. If students ask about these terms then discuss it outside of class and be mindful of students that haven't had a statistics class.}} | ||
}} | }} | ||
{{Definition | {{Definition | ||
|Error Bars | |Error Bars | ||
|A visual representation of the confidence interval on a graph. | |A visual representation of the confidence interval on a graph. | ||
|comments= | |comments= | ||
{{BoxTip | |||
|Error bars are typically drawn around a data point which is often the middle of the confidence interval. | |Error bars are typically drawn around a data point which is often the middle of the confidence interval. | ||
}} | }} | ||
}} | |||
{{Definition | {{Definition | ||
|Actively Open-minded Thinking (AOT)|A thinking style which emphasizes good reasoning independently of one's own beliefs by looking at issues from multiple perspectives, actively searching out ideas on both sides. It predicts more accurate calibration as well as the ability to evaluate argument quality objectively. | |Actively Open-minded Thinking (AOT)|A thinking style which emphasizes good reasoning independently of one's own beliefs by looking at issues from multiple perspectives, actively searching out ideas on both sides. It predicts more accurate calibration as well as the ability to evaluate argument quality objectively. | ||
}} | }} | ||
{{Definition | {{Definition | ||
|Growth Mindset | |Growth Mindset | ||
|A mindset where people believe "intelligence can be developed" and their abilities can be enhanced through learning. This ''also'' predicts more accurate calibration of credence levels. | |A mindset where people believe "intelligence can be developed" and their abilities can be enhanced through learning. This ''also'' predicts more accurate calibration of credence levels. | ||
}} | }} | ||
|-|Examples= | |-|Examples= | ||
<!-- Examples must be written with the Example template. --> | <!-- Examples must be written with the Example template. --> | ||
{{Example | {{Example | ||
|Lawyer Calibration | |Lawyer Calibration | ||
|[[File:Credence Level Lawyers.png|thumb|Lawyer calibration curve.]] | |[[File:Credence Level Lawyers.png|thumb|Lawyer calibration curve.]] | ||
Lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. [[:File:Insightful or Wishful- Lawyers' Ability to Predice Case Outcomes - Delahunty, Granhag, Hartwig, Loftus.pdf| | Lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. | ||
|links=[[:File:Insightful or Wishful- Lawyers' Ability to Predice Case Outcomes - Delahunty, Granhag, Hartwig, Loftus.pdf|Source]] | |||
}} | }} | ||
{{Example | {{Example | ||
|Nurse Calibration | |Nurse Calibration | ||
|[[File:Credence Level Nurses.png|thumb|Nurse calibration curve.]] | |[[File:Credence Level Nurses.png|thumb|Nurse calibration curve.]] | ||
Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. [[:File:Nurses' Risk Assessment Judgements; A Confidence Calibration Study - Yang, Thompson.pdf| | Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. | ||
|links=[[:File:Nurses' Risk Assessment Judgements; A Confidence Calibration Study - Yang, Thompson.pdf|Source]] | |||
|comments= | |comments= | ||
{{BoxCaution | |||
|title=Note | |title=Note | ||
|The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will ''not'' rain tomorrow with 70% confidence.) | |The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will ''not'' rain tomorrow with 70% confidence.) | ||
}} | }} | ||
}} | |||
{{Example | {{Example | ||
|Weather Forecaster Calibration | |Weather Forecaster Calibration | ||
|It is possible (and expected) for singular election forecasts to turn out to be wrong some of the time. It is fine as long as the predictions within that credence range are overall well calibrated. It is also impossible to calibrate the credence of a singular prediction, it is only possible for an aggregate of predictions. [[:File:The Polls Weren't Great. But That's Pretty Normal. - Silver.pdf| | |It is possible (and expected) for singular election forecasts to turn out to be wrong some of the time. It is fine as long as the predictions within that credence range are overall well calibrated. It is also impossible to calibrate the credence of a singular prediction, it is only possible for an aggregate of predictions. | ||
|links=[[:File:The Polls Weren't Great. But That's Pretty Normal. - Silver.pdf|Source]] | |||
}} | }} | ||
{{Example | {{Example | ||
|Vague Verbiage and Words of Estimative Probability | |Vague Verbiage and Words of Estimative Probability | ||
|Words of estimative probability are vague terms used by intelligence analysts to convey the likelihood of an event without explicitly stating the associated probability. | |Words of estimative probability are vague terms used by intelligence analysts to convey the likelihood of an event without explicitly stating the associated probability. | ||
|links= | |links= | ||
{{LinkCard | |||
|url=https://en.wikipedia.org/wiki/Words_of_estimative_probability | |url=https://en.wikipedia.org/wiki/Words_of_estimative_probability | ||
|title=Words of Estimative Probability | |title=Words of Estimative Probability | ||
|description=Wikipedia article on words of estimative probability. | |description=Wikipedia article on words of estimative probability. | ||
}} | }} | ||
}} | |||
{{Exemplary | |||
|{{Blockquote|Weather forecasters often seem wrong, but they only give probabilities, and their probabilities are really well calibrated. So we should trust weather forecasters, but remember that a 90% chance of rain also means a 10% chance of no rain.}} | |||
{{Blockquote|Being well calibrated does not require always predicting the correct outcome but requires being able to predict how many times one will be wrong.}} | |||
}} | |||
|-|Common Misconceptions= | |-|Common Misconceptions= | ||
<!-- Misconceptions must be written with the Misconception template. --> | <!-- Misconceptions must be written with the Misconception template. --> | ||
{{Misconception|He seems super confident, and she said she was only 85% sure, so we should trust him over her.|It is more important to have the self-awareness of how often they are wrong, than to always insist they are right. If possible, use the outcomes of their past predictions to evaluate the calibration of their confidence before placing trust in them.}} | {{Misconception|He seems super confident, and she said she was only 85% sure, so we should trust him over her.|It is more important to have the self-awareness of how often they are wrong, than to always insist they are right. If possible, use the outcomes of their past predictions to evaluate the calibration of their confidence before placing trust in them.}} | ||
{{Misconception|Am I well calibrated in this particular prediction?|The question doesn't really make sense. You can't have a calibration for a single prediction. A calibration level is only meaningfully defined over a collection of predictions for which you can see how many came ultimately true.}} | {{Misconception|Am I well calibrated in this particular prediction?|The question doesn't really make sense. You can't have a calibration for a single prediction. A calibration level is only meaningfully defined over a collection of predictions for which you can see how many came ultimately true.}} | ||
{{Misconception|Error bars say that the true value must lie within that range.|This is incorrect. The error bars always have some credence level associated with them that the value is within that range.}} | {{Misconception|Error bars say that the true value must lie within that range.|This is incorrect. The error bars always have some credence level associated with them that the value is within that range.}} | ||
|-|Expanded Learning Goals= | |||
After this lesson, students should | |||
# Attitudes | |||
## Be wary when high degrees of confidence are claimed. | |||
## Suspect something is amiss if no finding from a scientific community is ever retracted. | |||
# Concept Acquisition | |||
## '''Confidence Interval:''' A range within which a true value of interest lies with a specified probability. | |||
### Most commonly a 95% confidence interval, which means there is a 5% chance the true value lies outside the range specified. | |||
## '''Error Bars:''' Smaller bars on a graph that show the range of likely true values around the observed value, typically a 95% confidence interval, or the observed value ± the standard error or standard deviation. | |||
## Scientific culture at its best reinforces the importance of uncertainty by offering respect and career advancement to people on the basis of calibration as well as accuracy. In attaching the ego to calibration as well as accuracy, this discourages scientists from being overly attached to their ideas being "right," encouraging them to prioritize truth over having been right. | |||
## People (including many experts) tend to overestimate their accuracy at high confidence levels (and underestimate it at low confidence levels). | |||
## People often use a source's confidence as a cue to credibility, but appropriately discount confidence when they have evidence of poor calibration. | |||
# Concept Application | |||
## Compare the reliability of sources of information on the basis of their assiduousness in determining confidence levels and probabilistic ranges for results (e.g. confidence intervals, error bars). | |||
## Recognize that for scientific findings presented as having 95% certainty (for example), 5% of such results should be incorrect. | |||
## Not be fooled by criticisms of scientific communities for occasionally (~5% of the time, for example) having results later shown to be wrong. | |||
## Identify higher/lower accuracy and better/worse calibration in concrete examples. | |||
## Identify factors that lead to better calibration, and use these factors to predict and suggest ways to improve a person's calibration in a given scenario. | |||
</tabber> | </tabber> | ||
{{#restricted:{{Private:3.2 Calibration of Credence Levels}}}} | {{#restricted:{{Private:3.2 Calibration of Credence Levels}}}} | ||
{{NavCard|chapter=Lesson plans|text=All lesson plans|prev=3.1 Probabilistic Reasoning|next=4.1 Signal and Noise}} | {{NavCard|chapter=Lesson plans|text=All lesson plans|prev=3.1 Probabilistic Reasoning|next=4.1 Signal and Noise}} | ||
[[Category:Lesson plans]] | [[Category:Lesson plans]] | ||
Latest revision as of 22:26, 11 June 2026
How do we know what the appropriate degree of confidence for any statement of fact should be? Even experts often suffer from overconfidence. We introduce practical techniques to calibrate appropriate levels of confidence as well as psychological attitudes that motivate one to improve one's calibration.
The Lesson in Context
This lesson continues the previous lesson, 3.1 Probabilistic Reasoning, and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the Future Predictions activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration.
Takeaways
After this lesson, students should
- Be wary of high levels of confidence.
- Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars).
- Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration.
Confidence Interval
Confidence interval, when describing an instrumental measurement, corresponds to the statistical uncertainty.
When one presents a scientific measurement, they are usually not presenting a specific value. They're presenting a range that the value is within and the likelihood that the true value is within that range.
Different fields have different standards for confidence intervals. Physicists typically choose confidence intervals within which they are 68% sure the true value lies. Psychologists typically use 95%.
There are other related terms like "standard deviation," "[math]\displaystyle{ \sigma }[/math]," and "standard error." Try to avoid this jargon. If students ask about these terms then discuss it outside of class and be mindful of students that haven't had a statistics class.
Error Bars
Error bars are typically drawn around a data point which is often the middle of the confidence interval.
Actively Open-minded Thinking (AOT)
Growth Mindset
Lawyer Calibration

Lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up.
Nurse Calibration

Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time.
Note
The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will not rain tomorrow with 70% confidence.)
Weather Forecaster Calibration
Vague Verbiage and Words of Estimative Probability
Exemplary Quotes
“Weather forecasters often seem wrong, but they only give probabilities, and their probabilities are really well calibrated. So we should trust weather forecasters, but remember that a 90% chance of rain also means a 10% chance of no rain.”
“Being well calibrated does not require always predicting the correct outcome but requires being able to predict how many times one will be wrong.”
He seems super confident, and she said she was only 85% sure, so we should trust him over her.
Am I well calibrated in this particular prediction?
Error bars say that the true value must lie within that range.
After this lesson, students should
- Attitudes
- Be wary when high degrees of confidence are claimed.
- Suspect something is amiss if no finding from a scientific community is ever retracted.
- Concept Acquisition
- Confidence Interval: A range within which a true value of interest lies with a specified probability.
- Most commonly a 95% confidence interval, which means there is a 5% chance the true value lies outside the range specified.
- Error Bars: Smaller bars on a graph that show the range of likely true values around the observed value, typically a 95% confidence interval, or the observed value ± the standard error or standard deviation.
- Scientific culture at its best reinforces the importance of uncertainty by offering respect and career advancement to people on the basis of calibration as well as accuracy. In attaching the ego to calibration as well as accuracy, this discourages scientists from being overly attached to their ideas being "right," encouraging them to prioritize truth over having been right.
- People (including many experts) tend to overestimate their accuracy at high confidence levels (and underestimate it at low confidence levels).
- People often use a source's confidence as a cue to credibility, but appropriately discount confidence when they have evidence of poor calibration.
- Confidence Interval: A range within which a true value of interest lies with a specified probability.
- Concept Application
- Compare the reliability of sources of information on the basis of their assiduousness in determining confidence levels and probabilistic ranges for results (e.g. confidence intervals, error bars).
- Recognize that for scientific findings presented as having 95% certainty (for example), 5% of such results should be incorrect.
- Not be fooled by criticisms of scientific communities for occasionally (~5% of the time, for example) having results later shown to be wrong.
- Identify higher/lower accuracy and better/worse calibration in concrete examples.
- Identify factors that lead to better calibration, and use these factors to predict and suggest ways to improve a person's calibration in a given scenario.
Additional Content
You must be logged in to see this content.
