6.2 Hill's Criteria: Difference between revisions
More actions
// Edit via Wikitext Extension for VSCode |
Winstonyin (talk | contribs) No edit summary |
||
| Line 4: | Line 4: | ||
{{Navbox}} | {{Navbox}} | ||
== The Lesson in Context == | |||
<!-- Always begin section with a description of this lesson in relation to the course as a whole. --> | |||
This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples. | |||
<!-- Expandable section relating this lesson to earlier lessons. --> | |||
{{Expand|Relation to Earlier Lessons| | |||
{{ContextLesson|2.2 Systematic and Statistical Uncertainty}} | |||
{{ContextRelation|When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.}} | |||
}} | |||
{{ContextLesson|3.1 Probabilistic Reasoning}} | |||
{{ContextRelation|In the upcoming lesson on Probabilistic Reasoning, students will explore how to use quantified credence (i.e., confidence) levels to track and communicate their certainty about claims and predictions. For example, a .50 credence level indicates a 50-50 chance the claim is false or true, whereas a 1.00 credence level indicates complete certainty. Because it is difficult to impossible to conclusively infer causation without an RCT, Hill's criteria are typically used to infer causation with varying degrees of confidence, such that each additional piece of evidence (especially of new types) increases the confidence that the causal claim is true.}} | |||
{{ContextLesson|6.1 Correlation and Causation}} | |||
{{ContextRelation|RCTs are introduced as ideal experiments. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of a causation. | |||
Causation is defined as correlation under manipulation. When manual manipulation is not possible, Hill's criteria can still help establish causation.}} | |||
<!-- Expandable section relating this lesson to later lessons. --> | |||
{{Expand|Relation to Later Lessons| | |||
{{ContextLesson|11.2 When Is Science Suspect}} | |||
{{ContextRelation|In this upcoming lesson, students will explore how science has sometimes been used to justify the oppression of certain human groups, e.g. by inferring from current differences in achievement that there are fundamental differences in capacity, when in fact these differences in achievement can be fully explained by differences in opportunity. Because it is impossible to randomly assign individuals to human groups such as "African-Americans" or "women," all of this science is observational, and subject to the limitations of observational research. In evaluating non-RCT evidence, it is important to remember these limitations of observational as opposed to experimental research, especially when data may be used to exacerbate existing injustices.}} | |||
}} | |||
== Useful Links == | == Useful Links == | ||
| Line 65: | Line 98: | ||
== Context == | == Context == | ||
== Recommended Outline == | == Recommended Outline == | ||
Revision as of 07:28, 2 August 2023

Building on Correlation and Causation, we examine how to collect evidence for causality in more difficult cases.
The Lesson in Context
This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.
- In the upcoming lesson on Probabilistic Reasoning, students will explore how to use quantified credence (i.e., confidence) levels to track and communicate their certainty about claims and predictions. For example, a .50 credence level indicates a 50-50 chance the claim is false or true, whereas a 1.00 credence level indicates complete certainty. Because it is difficult to impossible to conclusively infer causation without an RCT, Hill's criteria are typically used to infer causation with varying degrees of confidence, such that each additional piece of evidence (especially of new types) increases the confidence that the causal claim is true.
- RCTs are introduced as ideal experiments. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of a causation.
Causation is defined as correlation under manipulation. When manual manipulation is not possible, Hill's criteria can still help establish causation.
Useful Links
- Causation Paper 2
- Causation Paper 3
- Three Column Overview of the Week
- Lesson Slides (2023 Master)
- Website Page
Readings and Assignments
Lecture Video
Learning Goals
After this lesson, students should
- Identify cases in which "ideal" RCT experiments are not possible, due to ethical or practical constraints.
- For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
- Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and consistency across contexts.
- Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.
Definitions
Hill's criteria:
- Prior plausibility
- Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
- Temporality/Temporal Sequence
- Did the hypothesized cause precede the effect?
- Specificity
- Specific predictions for specific consequences that have come true are less likely to be caused by other factors.
- Dose-response Curve
- Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?
- Consistency Across Contexts
- Does the correlation appear across diverse contexts?
| It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off). |
Examples
- Scrotal cancer in chimney sweeps, example by John.
- The causal connection between leaded gasoline across the world and violent crimes in many countries. (See video above, 17:57.)
- Prior Plausibility: High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
- Temporality: In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
- Specificity: Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
- Dose-response Curve: See temporality.
- Consistency Across Contexts: The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.
Common Misconceptions
- Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.
- A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.
- Students are quick to notice small sample size, slower to notice problems with experimental design.
- Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.
- Identifying natural experiments is also difficult, probably because many students find the principles of RCTs slippery.
Context
Recommended Outline
Before Class
- Review PlayPosit and discussion questions and ask faculty, Gabriel, or Emlen any questions you have.
During Class
| 5 Minutes | Come up with some fun way to assign the roles of spokesperson and notetaker (e.g. earliest birthday in the year, lives furthest from campus). Remind them of the responsibilities of these roles. |
| 15 Minutes | First set of discussion questions: Concept review. |
| 25 Minutes | #Paper Analysis (Part 2): #Vietnam War Draft Lottery and #Sunset Time and Social Jetlag |
| 15 Minutes | Further discussions: How to apply Hill's criteria to the social jetlag study. |
| 15 Minutes | Go over the COVID-19 example. |
| 5 Minutes | Answer any lingering questions. These activities have a tendency to go long, so some extra wiggle room is left at the end. |
Lesson Content
Concept Review
- (5 min) Review: What are the three essential elements of a Randomized Controlled Trial?
- Random Assignment
- Control Group/Condition
- Trial/Experimental Intervention
Note that random sampling is not essential to the validity of an RCT. It merely affects the generalizability of the results of the RCT.
- For a child, does spending more than 22 hours in their home per day over their childhood increase their risk of developing myopia?
It is not hard to imagine an ideal RCT for this, but it is unethical/impractical to perform the intervention. - On the level of individual towns, does implementing universal basic income reduce the rate of violent crimes in 10 years? In 150 years?
It is conceivable to do an RCT over 10 years, albeit expensive. However, it is hard to do this experiment on a large enough sample. It is also much less practical to track such an experiment over 150 years. It is difficult to randomise the towns.
Paper Analysis (Part 2)
(25 min)
This is a continuation of last lesson's paper analysis, in which we looked at a classic example of a proper RCT. Here, we will look at two more papers in which an RCT is difficult to conduct, but a strong case can still be made about causation based on experimental results.
| Strictly speaking, the experiments in the following two papers are not RCTs, since the experimenter never performed any intervention. However, their causal conclusions are nearly as strong as an RCT. Hill's criteria, on the other hand, are useful when even this kind of experiment is impossible to do. We will go through an example of Hill's criteria in #Covid-19 below. |
For each of the two papers, students should try to answer the following.
- What causal relationship is this paper trying to study? What is the hypothesis?
- What, if it exists at all, is the experimental intervention? (i.e., What is the independent variable that is being manipulated)?
- What is the dependent variable that is being measured (i.e., the variable the researchers anticipate may be affected by the experimental intervention)?
- How is the independent variable manipulated? Are there control and intervention groups? Are those randomly assigned? Is the control condition a good one (only the independent variable is different, with all else kept equal)? Is this an RCT?
- What is the result of the experiment?
- Can the causal relationship in question 1 be concluded from the experimental results? If not, what, if anything, can be concluded? How confident are you in this conclusion?
- Can you think of an alternative explanation for the data?
Vietnam War Draft Lottery
Example of natural experiment, where the intervention and assignment are not performed by the experimenter but by environmental factors: Lifetime Earnings and the Vietnam Era Draft Lottery
|
Upshot: This is a natural experiment. All the aspects of an RCT are still satisfied, but the randomized selection and intervention process is not performed by the experimenter but by an external agent/natural process.
Sunset Time and Social Jetlag
Not technically an RCT but could be argued to have comparable strength: Sunset Time and the Economic Effects of Social Jetlag
|
Discussion Questions
In your small group, explain Hill's criteria using aspects of the Sunset Time and Social Jetlag study as examples.
|
Covid-19
(15 min)
There has been a lot of controversy regarding mask wearing policies to combat the spread of COVID-19. Suppose it were very difficult to conduct a randomized controlled trial to study the effect of community mask wearing on COVID spread. (Such RCTs have in fact been conducted; see this paper.) However, we'd still like to find out whether there is such a causal connection using Hill's criteria.
Work in small groups. For each of the Hill's criteria, what evidence, if observed, would support that criterion?
- Plausible mechanism
Masks physically blocks one's mouth from emitting/receiving virus-ridden aerosols. - Temporal sequence
Do COVID cases reduce after a mask mandate has been instituted in an area/business/school? - Consistency across contexts
Is this reduction in cases after a mask mandate observed in many countries/areas/businesses/schools? - Dose-response curve
Does the extent of decrease in COVID cases correlate with the percentage of mask wearers in an area? - Specificity
Is it the case that counties/businesses/schools that implemented a mask mandate have a dramatically reduced case rate compared to those that did not?