3.2 Calibration of Credence Levels: Difference between revisions
More actions
// via Wikitext Extension for VSCode |
// via Wikitext Extension for VSCode |
||
| Line 64: | Line 64: | ||
|[[File:Credence Level Nurses.png|thumb|Nurse calibration curve.]] | |[[File:Credence Level Nurses.png|thumb|Nurse calibration curve.]] | ||
Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. [[:File:Nurses' Risk Assessment Judgements; A Confidence Calibration Study - Yang, Thompson.pdf|(Source)]] | Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. [[:File:Nurses' Risk Assessment Judgements; A Confidence Calibration Study - Yang, Thompson.pdf|(Source)]] | ||
|comments={{BoxCaution|title=Note|The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will ''not'' rain tomorrow with 70% confidence.)}} | |comments= | ||
{{BoxCaution | |||
|title=Note | |||
|The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will ''not'' rain tomorrow with 70% confidence.) | |||
}} | |||
}} | }} | ||
Revision as of 15:32, 21 December 2024
How do we know what the appropriate degree of confidence for any statement of fact should be? Even experts often suffer from overconfidence. We introduce practical techniques to calibrate appropriate levels of confidence as well as psychological attitudes that motivate one to improve one's calibration.
The Lesson in Context
This lesson continues the previous lesson, 3.1 Probabilistic Reasoning, and addresses the importance of the calibration of credence levels. We illustrate with some real-life examples that people, even experts, often exhibit overconfidence. By resolving the Future Predictions activity from the previous lesson, we show the students how to improve their own calibration. Finally, with the Actively Open-minded Thinking and Growth Mindset surveys, we introduce psychological attitudes that can help motivate one to improve their calibration.
Takeaways
After this lesson, students should
- Be wary of high levels of confidence.
- Understand the different ways in which scientists in different fields discuss credence levels (e.g. 95% confidence interval, error bars).
- Appreciate that one can improve on the calibration of their credence levels, and one should strive to reach an accurate calibration.
Confidence Interval
Confidence interval, when describing an instrumental measurement, corresponds to the statistical uncertainty.
When one presents a scientific measurement, they are usually not presenting a specific value. They're presenting a range that the value is within and the likelihood that the true value is within that range.
Different fields have different standards for confidence intervals. Physicists typically choose confidence intervals within which they are 68% sure the true value lies. Psychologists typically use 95%.
There are other related terms like "standard deviation," "[math]\displaystyle{ \sigma }[/math]," and "standard error." Try to avoid this jargon. If students ask about these terms then discuss it outside of class and be mindful of students that haven't had a statistics class.
Error Bars
Error bars are typically drawn around a data point which is often the middle of the confidence interval.
Actively Open-minded Thinking (AOT)
Growth Mindset
Lawyer Calibration

Lawyers have a wide range of predictions (20%-100%) for how likely they are to win any given case. However, the actual results show a much narrower band (40%-70%). Really, the outcome is more of a toss-up. (Source)
Nurse Calibration

Calibration curves are shown for experienced nurses as well as students training to become nurses. In both cases, they appear to be overconfident when outcomes are more likely and underconfident when outcomes are less likely. The nurses did not seem to become better calibrated with time. (Source)
Note
The confidence values are all greater than 50%. Any prediction whose confidence is less than 50% can be rephrased as a prediction for the opposite with a confidence greater than 50%. (e.g. I'm 30% confident it will rain tomorrow means I think it will not rain tomorrow with 70% confidence.)
Weather Forecaster Calibration
Vague Verbiage and Words of Estimative Probability
He seems super confident, and she said she was only 85% sure, so we should trust him over her.
Am I well calibrated in this particular prediction?
Error bars say that the true value must lie within that range.
Additional Content
You must be logged in to see this content.
