Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

6.2 Hill's Criteria: Difference between revisions

From Sense & Sensibility & Science
// via Wikitext Extension for VSCode
 
(88 intermediate revisions by 3 users not shown)
Line 1: Line 1:
{{Navbox}}
{{Cover|6.2 Hill's Criteria}}


== Learning Goals ==
In the messy real world, an ideal randomized controlled trial may not always be feasible, for ethical or practical reasons. Even so, it is still possible to present compelling evidence for causation by considering whether the observed data satisfy a set of intuitive criteria introduced by Bradford Hill.


[Link to PlayPosit]
== The Lesson in Context ==


[Link to instructional video]
<!-- Always begin section with a description of this lesson in relation to the course as a whole. -->
This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.


[https://sensesensibilityscience.berkeley.edu/topic/7 More details]
<!-- Expandable section relating this lesson to other lessons. -->
{{Expand|Relation to Other Lessons|
'''Earlier Lessons'''
{{ContextLesson|2.2 Systematic and Statistical Uncertainty}}
{{ContextRelation|When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.}}
{{ContextLesson|3.1 Probabilistic Reasoning}}
{{ContextRelation|A credence level is always associated with any scientific claim of causation. As each Hill's criterion adds to a case for causation, so our credence level for a causal relation increases. However, unlike in the case of RCTs, it may be difficult to quantify.}}
{{ContextLesson|6.1 Correlation and Causation}}
{{ContextRelation|RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.}}
{{Line}}
'''Later Lessons'''
{{ContextLesson|11.2 When Is Science Suspect}}
{{ContextRelation|Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.}}
}}
== Takeaways ==
 
<tabber>
 
|-|Learning Goals=


After this lesson, students should
After this lesson, students should
# Identify cases in which “ideal” experiments are not possible, due to ethical or practical constraints.  
<!-- Learning goals are written as a numbered list. -->
# Identify cases in which "ideal" RCT experiments are not possible, due to ethical or practical constraints.  
# For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.   
# For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.   
# Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and multiple consistency across contexts.
# Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and consistency across contexts.
# Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.
# Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.


=== Definitions ===
|-|Definitions=
<!-- Definitions must be written with the Definition and Subdefinition templates. The first Definition should have the "first=yes" flag at the end. -->
{{Definition|Hill's criteria|A group of criteria that suggest possible causation even in the absence of an RCT.|first=yes}}
{{Subdefinition|Prior Plausibility|Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?}}
{{Subdefinition|Temporality/Temporal Sequence|Did the hypothesized cause precede the effect?}}
{{Subdefinition|Specificity|Specific predictions for specific consequences that have come true are less likely to be caused by other factors.}}
{{Subdefinition|Dose-response Curve|Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?}}
{{Subdefinition|Consistency Across Contexts|Does the correlation appear across diverse contexts?}}
{{BoxCaution|This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.}}
{{BoxCaution|It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).}}
<br />


Hill's criteria:
|-|Examples=
* '''Prior plausibility'''
*: Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
* '''Temporality/temporal sequence'''
*: Did the hypothesized cause precede the effect?
* '''Specificity'''
*: Specific predictions for specific consequences that have come true are less likely to be caused by other factors.
* '''Dose-response curve'''
*: Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?
* '''Consistency across contexts'''
*: Does the correlation appear across diverse contexts?
{{Caution|It is not necessary for all of Hill’s criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).}}


=== Common Misconceptions ===
{{Example
|Leaded Gasoline and Violent Crimes
|The causal connection between leaded gasoline across the world and violent crimes in many countries.
* '''Prior Plausibility''': High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
* '''Temporality''': In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
* '''Specificity''': Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
* '''Dose-response Curve''': See temporality.
* '''Consistency Across Contexts''': The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.
|links={{LinkCard
|url=https://youtu.be/IV3dnLzthDA?t=1077
|title=Leaded Gasoline and Violent Crime
|description=Video on the causal connection between leaded gasoline and violent crime.}}
}}
{{Exemplary
|{{Blockquote|Ok, I agree that 'correlation doesn't prove causation' in general, but in a case like this where we have lots of other kinds of evidence it sure gives us a pretty strong guess about causation.}}
{{Blockquote|About a third of people admitted to relationship cycling, with similar rates among couples of different sexual orientations. That behavior, the researchers found, was correlated with increases in psychological distress, even after accounting for other factors that can influence mental health, such as demographic information, marital and family status, sexual orientation and related stressors. The more on-off cycles a person reported, Monk says, the larger the increases in depression and anxiety seemed to be.|[http://time.com/5377841/off-on-relationships/ Source]}}
{{Blockquote|From this, it follows that predicting future landscape patterns is difficult because contingencies may be unanticipated or even unpredictable. When similar locations can arise from different histories, and similar histories can produce different outcomes (e.g., Ernoult et al. 2006), it is not easy to infer causation.|''Landscape Ecology in Theory and Practice'', pg. 58}}
{{Blockquote|One hundred years after that, French chemist Antoine Lavoisier used a device called an "ice calorimeter" to gauge the energy burn from animals —[https://library.missouri.edu/exhibits/food/lavoisier.html like guinea pigs]— in cages by watching how quickly ice or snow around the cages melted. This research suggested that the heat and gases respired by animals, including humans, related to the energy they burn.|[https://www.vox.com/2018/9/4/17486110/metabolism-diet-fast-weight-loss Source]}}
{{Blockquote|I don't think there is a simple answer, because it is not clear that there is an increased risk of premature death - these are all observational studies and we are not controlling the amount of alcohol people consume and then analyzing the risk of death.|[https://www.reuters.com/article/us-health-alcohol/even-one-drink-a-day-linked-to-lower-life-expectancy-idUSKBN1I42H6 Dr. Eugene Yang]}}
}}


* ''Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.''
|-|Common Misconceptions=
*: A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.
* Students are quick to notice small sample size, slower to notice problems with experimental design.
* Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.
* Identifying natural experiments is also difficult, probably because many students find the principles of RCTs slippery.


== Context ==
<!-- Misconceptions must be written with the Misconception template. The first Misconception should have the "first=yes" flag at the end. -->
{{Misconception|Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.|A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.|first=yes}}
{{Misconception|Students are quick to notice small sample size, slower to notice problems with experimental design.|Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.}}


This is a discussion-based lesson that familiarises students with the concept of Hill’s criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble remembering their names, and they would benefit from a diverse range of examples.
|-|Expanded Learning Goals=


=== Before ===
After this lesson, students should
: '''[[2.2 Causation and Correlation]]'''
# Attitudes
:* RCTs are introduced as ideal experiments. These are often not possible due to resource or ethical considerations, and we must resort to Hill’s criteria to examine the plausibility of a causation.
## Appreciate that we can sometimes get very good evidence for a causal hypothesis, even in the absence of decisive RCTs.
:* Causation is defined as correlation under manipulation. When manual manipulation is not possible, Hill’s criteria can still help establish causation.
## Be wary of potential confounds in apparent evidence for causality.
: '''[[3.1 Systematic and Statistical Uncertainty]]'''
# Concept Acquisition
:: When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.
## In many cases it is not possible to conduct a true RCT to test causality, for practical or ethical reasons.
 
## There are non-RCT forms of evidence for causal hypotheses, which are less conclusive than RCTs but together can offer strong evidence for causation. These include:
=== After ===
### '''Prior plausibility:''' Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
 
### '''Temporality/temporal sequence:''' Did the hypothesized cause precede the effect?
: '''[[6.1 Probabilistic Thinking]]'''
### '''Dose-response curve:''' Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect?
:: In the upcoming lesson on Probabilistic Reasoning, students will explore how to use quantified credence (i.e., confidence) levels to track and communicate their certainty about claims and predictions. For example, a .50 credence level indicates a 50-50 chance the claim is false or true, whereas a 1.00 credence level indicates complete certainty. Because it is difficult to impossible to conclusively infer causation without an RCT, Hill’s criteria are typically used to infer causation with varying degrees of confidence, such that each additional piece of evidence (especially of new types) increases the confidence that the causal claim is true.
### '''Consistency across contexts:''' Does the correlation appear across diverse contexts?
: '''[[9.2 When Is Science Suspect]]'''
# Concept Application
:: In this upcoming lesson, students will explore how science has sometimes been used to justify the oppression of certain human groups, e.g. by inferring from current differences in achievement fundamental differences in capacity, when in fact these differences in achievement can be fully explained by differences in opportunity. Because it is impossible to randomly assign individuals to human groups such as “African-Americans” or “women,” all of this science is observational, and subject to the limitations of observational research. In evaluating non-RCT evidence, it is important to remember these limitations of observational as opposed to experimental research, especially when data may be used to exacerbate existing injustices.
## For a given causal hypothesis and imperfect study, identify the imperfections (e.g., sample size, lack of randomization, lack of control) and explain how these imperfections impact claims of causality.
 
## Identify potential confounds in RCT and non-RCT studies.
== Recommended Outline ==
## Sketch out methods for eliminating potential confounds in sample RCT or non-RCT studies.
 
## For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
=== Before Class ===
## For a given scenario in which a causal hypothesis is being made, describe an ideal experiment/set of experiments to test the hypothesis and rule out alternative hypotheses.
 
## Identify cases in which 'ideal' experiments are not possible, due to ethical or practical constraints.
* [Any essential logistical things that need to be done for this class]
## Evaluate the strength of causal claims when various sources of evidence are used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
* Prepare a seating chart.
## Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
* Review PlayPosit and discussion questions and ask faculty, Gabriel, or Emlen any questions you have.
</tabber>
* (Optional) Prepare a presentation.
 
=== During Class ===
 
* (5 min) Come up with some fun way to assign the roles of spokesperson and notetaker (e.g. earliest birthday in the year, lives furthest from campus). Remind them of the responsibilities of these roles.
* (21 min) [[#Concept Review|First set of discussion questions]]: Concept review & time zone effect on sleep.
* (16 min) [[#Chinese Restaurant Syndrome|Second set of discussion questions]]: MSG and headaches.
* (13 min) [[#Covid-19|Third set of discussion questions]]: Covid-19.
* (20 min) [[#Causal Link Brainstorming|Fourth set of discussion questions]]: Broader questions about evidence for causality.
* (5 min) Collect questions for plenary.
 
== Lesson Content ==
 
=== Clicker Question ===
 
# Question
## Option 1
## Options 2
 
=== Activity 1: Name ===
 
[Brief description of and motivation for the activity]
{{Caution|Common misconceptions and any useful tricks, tips, guidelines, or other background}}
 
==== Instructions ====
 
==== Discussion Questions ====
 
# Question 1
## Subquestion a {{Caution|Possible misconception that may need to be corrected and clarified.|small=right}}
##: {{Answer|Intended answer to the above question.}}
## Subquestion b
 
== Collect Questions for Plenary ==
 
(5 min) Collect remaining questions from the students for faculty in plenary (can be questions for clarification, extension, discussion, etc.), and add [ here].


[[Category:Lesson Plans]]
{{#restricted:{{Private:6.2 Hill's Criteria}}}}
{{NavCard|chapter=Lesson plans|text=All lesson plans|prev=6.1 Correlation and Causation|next=7.1 Causation, Blame, and Policy}}
[[Category:Lesson plans]]

Latest revision as of 23:00, 11 June 2026

In the messy real world, an ideal randomized controlled trial may not always be feasible, for ethical or practical reasons. Even so, it is still possible to present compelling evidence for causation by considering whether the observed data satisfy a set of intuitive criteria introduced by Bradford Hill.

The Lesson in Context

This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.

Earlier Lessons

2.2 Systematic and Statistical Uncertainty
  • When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.
3.1 Probabilistic Reasoning
  • A credence level is always associated with any scientific claim of causation. As each Hill's criterion adds to a case for causation, so our credence level for a causal relation increases. However, unlike in the case of RCTs, it may be difficult to quantify.
6.1 Correlation and Causation
  • RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.

Later Lessons

11.2 When Is Science Suspect
  • Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.

Takeaways

After this lesson, students should

  1. Identify cases in which "ideal" RCT experiments are not possible, due to ethical or practical constraints.
  2. For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
  3. Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and consistency across contexts.
  4. Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.

Hill's criteria

A group of criteria that suggest possible causation even in the absence of an RCT.
  • Prior Plausibility
Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
  • Temporality/Temporal Sequence
Did the hypothesized cause precede the effect?
  • Specificity
Specific predictions for specific consequences that have come true are less likely to be caused by other factors.
  • Dose-response Curve
Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?
  • Consistency Across Contexts
Does the correlation appear across diverse contexts?

This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.

It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).


Leaded Gasoline and Violent Crimes

The causal connection between leaded gasoline across the world and violent crimes in many countries.
  • Prior Plausibility: High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
  • Temporality: In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
  • Specificity: Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
  • Dose-response Curve: See temporality.
  • Consistency Across Contexts: The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.

Exemplary Quotes

Ok, I agree that 'correlation doesn't prove causation' in general, but in a case like this where we have lots of other kinds of evidence it sure gives us a pretty strong guess about causation.

About a third of people admitted to relationship cycling, with similar rates among couples of different sexual orientations. That behavior, the researchers found, was correlated with increases in psychological distress, even after accounting for other factors that can influence mental health, such as demographic information, marital and family status, sexual orientation and related stressors. The more on-off cycles a person reported, Monk says, the larger the increases in depression and anxiety seemed to be.

From this, it follows that predicting future landscape patterns is difficult because contingencies may be unanticipated or even unpredictable. When similar locations can arise from different histories, and similar histories can produce different outcomes (e.g., Ernoult et al. 2006), it is not easy to infer causation.

Landscape Ecology in Theory and Practice, pg. 58

One hundred years after that, French chemist Antoine Lavoisier used a device called an "ice calorimeter" to gauge the energy burn from animals —like guinea pigs— in cages by watching how quickly ice or snow around the cages melted. This research suggested that the heat and gases respired by animals, including humans, related to the energy they burn.

I don't think there is a simple answer, because it is not clear that there is an increased risk of premature death - these are all observational studies and we are not controlling the amount of alcohol people consume and then analyzing the risk of death.

Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.

A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.

Students are quick to notice small sample size, slower to notice problems with experimental design.

Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.

After this lesson, students should

  1. Attitudes
    1. Appreciate that we can sometimes get very good evidence for a causal hypothesis, even in the absence of decisive RCTs.
    2. Be wary of potential confounds in apparent evidence for causality.
  2. Concept Acquisition
    1. In many cases it is not possible to conduct a true RCT to test causality, for practical or ethical reasons.
    2. There are non-RCT forms of evidence for causal hypotheses, which are less conclusive than RCTs but together can offer strong evidence for causation. These include:
      1. Prior plausibility: Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
      2. Temporality/temporal sequence: Did the hypothesized cause precede the effect?
      3. Dose-response curve: Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect?
      4. Consistency across contexts: Does the correlation appear across diverse contexts?
  3. Concept Application
    1. For a given causal hypothesis and imperfect study, identify the imperfections (e.g., sample size, lack of randomization, lack of control) and explain how these imperfections impact claims of causality.
    2. Identify potential confounds in RCT and non-RCT studies.
    3. Sketch out methods for eliminating potential confounds in sample RCT or non-RCT studies.
    4. For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
    5. For a given scenario in which a causal hypothesis is being made, describe an ideal experiment/set of experiments to test the hypothesis and rule out alternative hypotheses.
    6. Identify cases in which 'ideal' experiments are not possible, due to ethical or practical constraints.
    7. Evaluate the strength of causal claims when various sources of evidence are used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
    8. Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.

Additional Content

You must be logged in to see this content.