3.2 Calibration of Credence Levels: Difference between revisions
More actions
No edit summary |
// Edit via Wikitext Extension for VSCode |
||
| Line 10: | Line 10: | ||
This lesson continues the previous lesson, [[3.1 Probabilistic Reasoning]], and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the [[3.1 Probabilistic Reasoning#Future Predictions|Future Predictions]] activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration. | This lesson continues the previous lesson, [[3.1 Probabilistic Reasoning]], and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the [[3.1 Probabilistic Reasoning#Future Predictions|Future Predictions]] activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration. | ||
<!-- Expandable section relating this lesson to | <!-- Expandable section relating this lesson to other lessons. --> | ||
{{Expand|Relation to Earlier Lessons | {{Expand|Relation to Other Lessons| | ||
'''Earlier Lessons''' | |||
{{ContextLesson|1.2 Shared Reality and Modeling}} | {{ContextLesson|1.2 Shared Reality and Modeling}} | ||
{{ContextRelation|Raft vs. pyramid: Having concrete methods to assess and calibrate credence levels allows scientists to comfortably discuss their imperfect results without being overly invested in ''being right''. This also leaves open the possibility that scientific claims can be revised or overturned in light of new evidence. This is a strength of the scientific method, rather than a weakness.}} | {{ContextRelation|Raft vs. pyramid: Having concrete methods to assess and calibrate credence levels allows scientists to comfortably discuss their imperfect results without being overly invested in ''being right''. This also leaves open the possibility that scientific claims can be revised or overturned in light of new evidence. This is a strength of the scientific method, rather than a weakness.}} | ||
{{ContextLesson|3.1 Probabilistic Reasoning}} | {{ContextLesson|3.1 Probabilistic Reasoning}} | ||
{{ContextRelation|The current lesson refines the previous one by introducing concrete ways to assess one's credence calibration and to improve upon them.}} | {{ContextRelation|The current lesson refines the previous one by introducing concrete ways to assess one's credence calibration and to improve upon them.}} | ||
}} | {{Line}} | ||
'''Later Lessons''' | |||
{{ContextLesson|5.1 False Positives and Negatives}} | {{ContextLesson|5.1 False Positives and Negatives}} | ||
{{ContextRelation|With accurately calibrated credence levels, decisions can be made on whether to act upon a prediction by considering the seriousness of its consequences. For example, a city should invest heavily in storm protection even when there is a 10% confidence that a strong hurricane will hit the city.}} | {{ContextRelation|With accurately calibrated credence levels, decisions can be made on whether to act upon a prediction by considering the seriousness of its consequences. For example, a city should invest heavily in storm protection even when there is a 10% confidence that a strong hurricane will hit the city.}} | ||
Revision as of 16:13, 30 August 2023

How do we know what the appropriate degree of confidence for any statement of fact should be? Even experts often suffer from overconfidence. We introduce practical techniques to calibrate appropriate levels of confidence as well as psychological attitudes that motivate one to improve one's calibration.
The Lesson in Context
This lesson continues the previous lesson, 3.1 Probabilistic Reasoning, and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the Future Predictions activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration.
Takeaways
After this lesson, students should
- Be wary of high levels of confidence.
- Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars).
- Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration.
Confidence Interval
Confidence interval, when describing an instrumental measurement, corresponds to the statistical uncertainty.
When one presents a scientific measurement, they are usually not presenting a specific value. They're presenting a range that the value is within and the likelihood that the true value is within that range.
Different fields have different standards for confidence intervals. Physicists typically choose confidence intervals within which they are 68% sure the true value lies. Psychologists typically use 95%.
There are other related terms like "standard deviation," "σ," and "standard error." Try to avoid this jargon. If students ask about these terms then discuss it outside of class and be mindful of students that haven't had a statistics class.
Error Bars
Error bars are typically drawn around a data point which is often the middle of the confidence interval.
Actively Open-minded Thinking (AOT)
Growth Mindset
Lawyer Calibration

- Lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. (Source)
Nurse Calibration

- Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. (Source)
Note
The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will not rain tomorrow with 70% confidence.)
Weather Forecaster Calibration
- It is possible (and expected) for singular election forecasts to turn out to be wrong some of the time. It is fine as long as the predictions within that credence range are overall well calibrated. It is also impossible to calibrate the credence of a singular prediction, it is only possible for an aggregate of predictions. (Source)
Vague Verbiage and Words of Estimative Probability
He seems super confident, and she said she was only 85% sure, so we should trust him over her.
Am I well calibrated in this particular prediction?
Error bars say that the true value must lie within that range.
Additional Content
You must be logged in to see this content.