Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

6.2 Hill's Criteria: Difference between revisions

From Sense & Sensibility & Science
// via Wikitext Extension for VSCode
 
(24 intermediate revisions by 2 users not shown)
Line 1: Line 1:
[[File:Topic Cover - 6.2 Hill's Criteria.png|thumb]]
{{Cover|6.2 Hill's Criteria}}


Building on Correlation and Causation, we examine how to collect evidence for causality in more difficult cases.
In the messy real world, an ideal randomized controlled trial may not always be feasible, for ethical or practical reasons. Even so, it is still possible to present compelling evidence for causation by considering whether the observed data satisfy a set of intuitive criteria introduced by Bradford Hill.
 
{{Navbox}}


== The Lesson in Context ==
== The Lesson in Context ==
Line 10: Line 8:
This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.
This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.


<!-- Expandable section relating this lesson to earlier lessons. -->
<!-- Expandable section relating this lesson to other lessons. -->
{{Expand|Relation to Earlier Lessons|
{{Expand|Relation to Other Lessons|
'''Earlier Lessons'''
{{ContextLesson|2.2 Systematic and Statistical Uncertainty}}
{{ContextLesson|2.2 Systematic and Statistical Uncertainty}}
{{ContextRelation|When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.}}
{{ContextRelation|When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.}}
Line 18: Line 17:
{{ContextLesson|6.1 Correlation and Causation}}
{{ContextLesson|6.1 Correlation and Causation}}
{{ContextRelation|RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.}}
{{ContextRelation|RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.}}
}}
{{Line}}
<!-- Expandable section relating this lesson to later lessons. -->
'''Later Lessons'''
{{Expand|Relation to Later Lessons|
{{ContextLesson|11.2 When Is Science Suspect}}
{{ContextLesson|11.2 When Is Science Suspect}}
{{ContextRelation|Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.}}
{{ContextRelation|Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.}}
}}
}}
== Takeaways ==
== Takeaways ==


Line 40: Line 37:
|-|Definitions=
|-|Definitions=
<!-- Definitions must be written with the Definition and Subdefinition templates. The first Definition should have the "first=yes" flag at the end. -->
<!-- Definitions must be written with the Definition and Subdefinition templates. The first Definition should have the "first=yes" flag at the end. -->
{{Definition|Hill's criteria||first=yes}}
{{Definition|Hill's criteria|A group of criteria that suggest possible causation even in the absence of an RCT.|first=yes}}
{{Subdefinition|Prior plausibility|Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?}}
{{Subdefinition|Prior Plausibility|Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?}}
{{Subdefinition|Temporality/Temporal Sequence|Did the hypothesized cause precede the effect?}}
{{Subdefinition|Temporality/Temporal Sequence|Did the hypothesized cause precede the effect?}}
{{Subdefinition|Specificity|Specific predictions for specific consequences that have come true are less likely to be caused by other factors.}}
{{Subdefinition|Specificity|Specific predictions for specific consequences that have come true are less likely to be caused by other factors.}}
Line 48: Line 45:
{{BoxCaution|This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.}}
{{BoxCaution|This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.}}
{{BoxCaution|It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).}}
{{BoxCaution|It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).}}
<br />


|-|Examples=
|-|Examples=


<!-- Example formatting is still experimental. -->
{{Example
* Scrotal cancer in chimney sweeps, example by John.
|Leaded Gasoline and Violent Crimes
<youtube>https://youtu.be/IV3dnLzthDA?t=1077</youtube>
|The causal connection between leaded gasoline across the world and violent crimes in many countries.
* The causal connection between leaded gasoline across the world and violent crimes in many countries. (See video above, 17:57.)
* '''Prior Plausibility''': High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
** '''Prior Plausibility''': High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
* '''Temporality''': In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
** '''Temporality''': In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
* '''Specificity''': Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
** '''Specificity''': Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
* '''Dose-response Curve''': See temporality.
** '''Dose-response Curve''': See temporality.
* '''Consistency Across Contexts''': The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.
** '''Consistency Across Contexts''': The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.
|links={{LinkCard
|url=https://youtu.be/IV3dnLzthDA?t=1077
|title=Leaded Gasoline and Violent Crime
|description=Video on the causal connection between leaded gasoline and violent crime.}}
}}
{{Exemplary
|{{Blockquote|Ok, I agree that 'correlation doesn't prove causation' in general, but in a case like this where we have lots of other kinds of evidence it sure gives us a pretty strong guess about causation.}}
{{Blockquote|About a third of people admitted to relationship cycling, with similar rates among couples of different sexual orientations. That behavior, the researchers found, was correlated with increases in psychological distress, even after accounting for other factors that can influence mental health, such as demographic information, marital and family status, sexual orientation and related stressors. The more on-off cycles a person reported, Monk says, the larger the increases in depression and anxiety seemed to be.|[http://time.com/5377841/off-on-relationships/ Source]}}
{{Blockquote|From this, it follows that predicting future landscape patterns is difficult because contingencies may be unanticipated or even unpredictable. When similar locations can arise from different histories, and similar histories can produce different outcomes (e.g., Ernoult et al. 2006), it is not easy to infer causation.|''Landscape Ecology in Theory and Practice'', pg. 58}}
{{Blockquote|One hundred years after that, French chemist Antoine Lavoisier used a device called an "ice calorimeter" to gauge the energy burn from animals —[https://library.missouri.edu/exhibits/food/lavoisier.html like guinea pigs]— in cages by watching how quickly ice or snow around the cages melted. This research suggested that the heat and gases respired by animals, including humans, related to the energy they burn.|[https://www.vox.com/2018/9/4/17486110/metabolism-diet-fast-weight-loss Source]}}
{{Blockquote|I don't think there is a simple answer, because it is not clear that there is an increased risk of premature death - these are all observational studies and we are not controlling the amount of alcohol people consume and then analyzing the risk of death.|[https://www.reuters.com/article/us-health-alcohol/even-one-drink-a-day-linked-to-lower-life-expectancy-idUSKBN1I42H6 Dr. Eugene Yang]}}
}}


|-|Common Misconceptions=
|-|Common Misconceptions=
Line 66: Line 75:
{{Misconception|Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.|A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.|first=yes}}
{{Misconception|Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.|A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.|first=yes}}
{{Misconception|Students are quick to notice small sample size, slower to notice problems with experimental design.|Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.}}
{{Misconception|Students are quick to notice small sample size, slower to notice problems with experimental design.|Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.}}
{{Misconception|Identifying natural experiments is also difficult, probably because many students find the principles of RCTs slippery.|}}


</tabber>
|-|Expanded Learning Goals=
 
== Useful Resources ==
 
<tabber>
 
|-|Lecture Video=
 
<br /><center><youtube>https://youtu.be/VYLznjXXXCU</youtube></center><br />
 
|-|Discussion Slides=
 
{{LinkCard
|url=https://docs.google.com/presentation/d/1sKDf-GTT_W0kfF8in0U3cTzNZpQJJgHacWgQdTAv8gs/edit?usp=drivesdk
|title=Discussion Slides Template
|description=The discussion slides for this lesson.
}}
<br />
 
|-|Handouts and Activities=
 
{{LinkCard
|url=https://sensibility.berkeley.edu/images/a/af/Angrist_-_2021_-_Draft_lottery_and_lifetime_earnings.pdf
|title=Lifetime Earnings and the Vietnam Era Draft Lottery
|description=The paper the students analyze in the first part of the paper analysis activity.}}
{{LinkCard
|url=https://sensibility.berkeley.edu/images/8/80/Giuntella_and_Mazzonna_-_2019_-_Sunset_time_and_social_jetlag.pdf
|title=Sunset time and the economic effects of social jetlag: evidence from US time zone borders
|description=The paper the students analyze in the second part of the paper analysis activity.}}
<br />
 
|-|Readings and Assignments=
 
{{LinkCard
|url=https://sensibility.berkeley.edu/images/b/b3/The_Environment_and_Disease_Association_or_Causation_-_Hill.pdf
|title=The Environment and Disease: Association or Causation?
|description={{Todo|write this}}
}}
<br />


After this lesson, students should
# Attitudes
## Appreciate that we can sometimes get very good evidence for a causal hypothesis, even in the absence of decisive RCTs.
## Be wary of potential confounds in apparent evidence for causality.
# Concept Acquisition
## In many cases it is not possible to conduct a true RCT to test causality, for practical or ethical reasons.
## There are non-RCT forms of evidence for causal hypotheses, which are less conclusive than RCTs but together can offer strong evidence for causation. These include:
### '''Prior plausibility:''' Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
### '''Temporality/temporal sequence:''' Did the hypothesized cause precede the effect?
### '''Dose-response curve:''' Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect?
### '''Consistency across contexts:''' Does the correlation appear across diverse contexts?
# Concept Application
## For a given causal hypothesis and imperfect study, identify the imperfections (e.g., sample size, lack of randomization, lack of control) and explain how these imperfections impact claims of causality.
## Identify potential confounds in RCT and non-RCT studies.
## Sketch out methods for eliminating potential confounds in sample RCT or non-RCT studies.
## For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
## For a given scenario in which a causal hypothesis is being made, describe an ideal experiment/set of experiments to test the hypothesis and rule out alternative hypotheses.
## Identify cases in which 'ideal' experiments are not possible, due to ethical or practical constraints.
## Evaluate the strength of causal claims when various sources of evidence are used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
## Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
</tabber>
</tabber>


 
{{#restricted:{{Private:6.2 Hill's Criteria}}}}
 
{{NavCard|chapter=Lesson plans|text=All lesson plans|prev=6.1 Correlation and Causation|next=7.1 Causation, Blame, and Policy}}
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
== Recommended Outline ==
 
=== Before Class ===
 
Review PlayPosit and discussion questions and ask faculty, Gabriel, or Emlen any questions you have.
 
=== During Class ===
 
{| class="wikitable" style="margin-left: 0px; margin-right: auto;"
|-
|5 Minutes
|Come up with some fun way to assign the roles of spokesperson and notetaker (e.g. earliest birthday in the year, lives furthest from campus). Remind them of the responsibilities of these roles.
|-
|15 Minutes
|[[#Concept Review|First set of discussion questions]]: Concept review.
|-
|25 Minutes
|[[#Paper Analysis (Part 2)]]: [[#Vietnam War Draft Lottery]] and [[#Sunset Time and Social Jetlag]]
|-
|15 Minutes
|[[#Discussion Questions|Further discussions]]: How to apply Hill's criteria to the social jetlag study.
|-
|15 Minutes
|Go over the [[#Covid-19|COVID-19 example]].
|-
|5 Minutes
|Answer any lingering questions. These activities have a tendency to go long, so some extra wiggle room is left at the end.
|}
 
== Lesson Content ==
 
=== Concept Review ===
 
# (5 min) Review: What are the three essential elements of a Randomized Controlled Trial? {{BoxAnswer|
## ''Random'' Assignment
## ''Control'' Group/Condition
## ''Trial''/Experimental Intervention}} {{BoxCaution|Note that random sampling is not essential to the validity of an RCT. It merely affects the generalizability of the results of the RCT.}}
# (10 min) For both of the following hypotheses, what RCT would you use to test it? Why would it be challenging to perform?
## For a child, does spending more than 22 hours in their home per day over their childhood increase their risk of developing myopia? {{BoxAnswer|title=Explanation|It is not hard to imagine an ideal RCT for this, but it is unethical/impractical to perform the intervention.}}
## On the level of individual towns, does implementing universal basic income reduce the rate of violent crimes in 10 years? In 150 years? {{BoxAnswer|It is conceivable to do an RCT over 10 years, albeit expensive. However, it is hard to do this experiment on a large enough sample. It is also much less practical to track such an experiment over 150 years. It is difficult to randomise the towns.}}
 
=== Paper Analysis (Part 2) ===
 
<!-- {{Todo|Respond to Emlen's concerns: 6.2 Hill's Criteria: Two studies on natural experiments in section, not really focused on hill's criteria. Perhaps better to do just one natural experiment, and more practice w/ Hill's. This is related to the funkiness of teaching RCTs before Hill's criteria. Switch out one natural experiment for something similar to the Covid & masks case, but where it's really impossible to do an RCT, maybe Chinese restaurant syndrome. Need a little more structure for the discussion of Hill's criteria on Sunset Time and Social Jetlag, i.e. instruct students explicitly to extract each of Hill's criteria from the studies. Concept Review currently RCTs, not Hill's criteria.}} -->
 
(25 min)
 
This is a continuation of [[6.1 Correlation and Causation#Paper Analysis (Part 1)|last lesson's paper analysis]], in which we looked at a classic example of a proper RCT. Here, we will look at two more papers in which an RCT is difficult to conduct, but a strong case can still be made about causation based on experimental results.
 
{{BoxCaution|Strictly speaking, the experiments in the following two papers are not RCTs, since the experimenter never performed any intervention. However, their causal conclusions are nearly as strong as an RCT. Hill's criteria, on the other hand, are useful when even this kind of experiment is impossible to do. We will go through an example of Hill's criteria in [[#Covid-19]] below.}}
 
For each of the two papers, students should try to answer the following.
# What causal relationship is this paper trying to study? What is the hypothesis?
# What, if it exists at all, is the experimental intervention? (i.e., What is the independent variable that is being manipulated)?
# What is the dependent variable that is being measured (i.e., the variable the researchers anticipate may be affected by the experimental intervention)?
# How is the independent variable manipulated? Are there control and intervention groups? Are those randomly assigned? Is the control condition a good one (only the independent variable is different, with all else kept equal)? Is this an RCT?
# What is the result of the experiment?
# Can the causal relationship in question 1 be concluded from the experimental results? If not, what, if anything, can be concluded? How confident are you in this conclusion?
# Can you think of an alternative explanation for the data?
 
==== Vietnam War Draft Lottery ====
 
Example of natural experiment, where the intervention and assignment are not performed by the experimenter but by environmental factors: [[:File:Angrist - 2021 - Draft lottery and lifetime earnings.pdf|Lifetime Earnings and the Vietnam Era Draft Lottery]]
{{BoxAnswer|
# Does being drafted to serve in the military cause one's income to change later in life?
# Whether one has served in the military (veteran status).
# One's lifetime income.
# The veteran status is manipulated by the Vietnam War draft lottery, which randomly selected young men to serve in the military. The control group consists of young men who were not drafted by the lottery. This seems like a good control, as the selection is supposedly not based on any factor that may introduce systematic bias to one's future income. There are other factors that allow people to get out of the draft. This paper ''attempts'' to account for this. This is a ''natural experiment''. {{Todo|@Gabriel write this more carefully.}}
# Veterans earn 15% less than non-veterans long after the draft and the war ended.
# Being drafted to serve in the military causes a long term reduction in one's income.
# Well-connected or wealthier young men may have had an easier time getting out of the draft, which is a possible confound. Unknown confounds are more likely in a natural experiment, where the differences between experimental group and control group are not as tightly controlled as in a deliberately created RCT.}}
 
Upshot: This is a natural experiment. All the aspects of an RCT are still satisfied, but the randomized selection and intervention process is not performed by the experimenter but by an external agent/natural process.
 
==== Sunset Time and Social Jetlag ====
 
Not technically an RCT but could be argued to have comparable strength: [[:File:Giuntella and Mazzonna - 2019 - Sunset time and social jetlag.pdf|Sunset Time and the Economic Effects of Social Jetlag]]
{{BoxAnswer|
# Does an extra hour of natural light in the evening (according to local time) cause a change in one's sleep duration?
# The time of sunset in one's local time zone.
# Sleep duration.
# People who live on either side of a time zone boundary are considered, since they live geographically close to each other but have a 1-hour difference in sunset time. We may treat those on the east side of a boundary to be the intervention group and those on the west side to be the control group. The people are certainly not randomly assigned to one group or the other, but it is presumed that there are no systematic differences in the populations on either side that would cause a big difference in sleep patterns, other than the time zones. This is not an RCT, but its strength in implying causal relationship may be argued to be comparable to that of an RCT.
# Those living in a time zone with one extra hour of daylight in the evening (an earlier sunset time) sleep 19 minutes less on average.
# Living in a time zone with one extra hour of daylight in the evening causes a 19-minute reduction in average sleep duration.
# This is a pretty convincing study, what with the huge representative Census sample and the arbitrary cutoff. However, it's possible that some existing geographic power/wealth differential enabled people one one side of the time zone split to set it at their advantage (e.g. people in cities), leaving the less powerful/wealthy on the darker side (e.g. people in more rural areas). If this occurred, the difference in sleep and health could be due to the pre-existing power or wealth differential, not the time zone.}}
 
==== Discussion Questions ====
 
In your small group, explain Hill's criteria using aspects of the ''Sunset Time and Social Jetlag'' study as examples. {{BoxAnswer|
* Hill's criteria is essentially a list of criteria that can add to the case of determining causality, when you can only observe correlation. The criteria are used when RCT's can't be performed. See [[#Definitions|definitions]] above.
* '''Prior plausibility''': The presence of sunlight and one's daily schedule are well known to affect sleep patterns. If the work schedule is shifted by one hour due to time zone, it should affect sleep patterns as well.
* '''Temporality/temporal sequence''': People have lived on one side of the time zone border before they developed this sleep pattern.
* '''Specificity''': The observed effect in sleep times is drastic across the time zone borders, and this effect is not observed elsewhere. This effect was also only observed in people that had to get up early for work.
* '''Dose-response curve''': The further away one lives from a time zone border (more accurate their sunset/work time is to the natural circadian rhythm), the less pronounced the effect on sleep times becomes. The effect is also absent among unemployed people. (Fig. 6)
* '''Consistency across contexts''': The effect on sleep times is observed across many different time zone borders scattered throughout the US.}}
 
=== Covid-19 ===
 
(15 min)
 
There has been a lot of controversy regarding mask wearing policies to combat the spread of COVID-19. Suppose it were very difficult to conduct a randomized controlled trial to study the effect of community mask wearing on COVID spread. (Such RCTs have in fact been conducted; see [https://www.science.org/doi/10.1126/science.abi9069 this paper].) However, we'd still like to find out whether there is such a causal connection using Hill's criteria.
 
Work in small groups. For each of the Hill's criteria, what evidence, if observed, would support that criterion?
# Plausible mechanism {{BoxAnswer|Masks physically blocks one's mouth from emitting/receiving virus-ridden aerosols.}}
# Temporal sequence {{BoxAnswer|Do COVID cases reduce ''after'' a mask mandate has been instituted in an area/business/school?}}
# Consistency across contexts {{BoxAnswer|Is this reduction in cases after a mask mandate observed in many countries/areas/businesses/schools?}}
# Dose-response curve {{BoxAnswer|Does the extent of decrease in COVID cases correlate with the percentage of mask wearers in an area?}}
# Specificity {{BoxAnswer|Is it the case that counties/businesses/schools that implemented a mask mandate have a dramatically reduced case rate compared to those that did not?}}
 
<!-- == Overflow ==
 
<div class="toccolours mw-collapsible mw-collapsed" style="overflow:auto;">
<div style="font-weight:bold;line-height:1.6;">Extra content that's not currently part of the official lesson plan.</div>
<div class="mw-collapsible-content">
 
=== Chinese Restaurant Syndrome ===
 
(This has been a good quiz question when it wasn't a discussion topic.)
 
Starting from the late 1960s, some people in the USA noticed that they had a headache after eating at Chinese restaurants. They attributed their symptoms to the consumption of monosodium glutamate (MSG), a common food additive in East Asian cuisines. "Chinese Restaurant Syndrome" was coined as a term referring to the negative health effects of MSG.
{{BoxCaution|"Chinese restaurant syndrome" was the popular term historically used and it highlights the underlying racism and the bias associated with the claims. We will be discussing prejudices and biases in society clouding our sound scientific judgment in a future topic.}}
# (4 min) Describe a randomized controlled trial (RCT) that can test the hypothesis that "consumption of MSG causes headaches." {{BoxAnswer|Randomly split the participants into two groups (random assignment). Give one group a MSG-filled meal (experimental condition). Give the other group a meal low in MSG (control condition). Observe if the participants experience a headache after the meal.}}
# (10 min) Suppose that for some reason a RCT could not be done, but after some research, the following facts are known.
#* MSG also appears in Western cuisine, naturally in tomatoes and artificially added in many processed foods, and Americans who reported Chinese Restaurant Syndrome did not experience such symptoms after eating a Western meal containing MSG.
#* People in East Asian countries, who consistently consume meals containing high amounts of MSG, do not report such symptoms.
#* No biological pathway is known that connects MSG consumption and headaches.
## (2 min) Why might a RCT not be able to be conducted? {{BoxAnswer|An RCT might not be able to be conducted for a variety of ethical or practical reasons.}}
## (6 min) Each of the three points above can be used as evidence against a causal relationship between MSG consumption and headaches. For each point, discuss which of Hill's criteria for causality is ''not'' satisfied as a result of that fact. {{BoxAnswer|(1) Consistency across different contexts. (2) Dose-response curve and consistency across contexts. (3) Plausible mechanism.}}
 
Based on Hill's Criteria we can say that there is not a general causation between MSG and the symptoms claimed by "Chinese Restaurant Syndrome." See the quote at [https://www.fda.gov/food/food-additives-petitions/questions-and-answers-monosodium-glutamate-msg this link]:
: "These adverse event reports helped trigger FDA to ask the independent scientific group Federation of American Societies for Experimental Biology (FASEB) to examine the safety of MSG in the 1990s. FASEB's report concluded that MSG is safe. The FASEB report identified some short-term, transient, and generally mild symptoms, such as headache, numbness, flushing, tingling, palpitations, and drowsiness that may occur in some sensitive individuals who consume 3 grams or more of MSG without food. However, a typical serving of a food with added MSG contains less than 0.5 grams of MSG. Consuming more than 3 grams of MSG without food at one time is unlikely."
 
=== Causal Link Brainstorming ===
 
# (5 min) Brainstorm three causal links that you are pretty confident are real without RCT evidence.
# (4 min) Why are you confident about these causal links? {{BoxCaution|small=right|Think about causal links in everyday life, like allergies or breaking a glass.}} {{BoxAnswer|There are any number of answers to this question (i.e. Rain clouds cause rain, The sun heats up objects, Exercise causes muscle growth, Heat will melt butter, Pressing the unlock button on your keys unlocks your car etc.)}}
# (5 min) Can you come up with a situation where someone might come to the wrong conclusions about these causal links? {{BoxAnswer|Temporality may cause you to draw the wrong conclusions, especially when it is immediate (i.e. In the PlayPosit John talked about the power outage and someone in a family argument and pounding the table and the lights go out; Homeopathic remedies can lead to the belief that the remedy caused the healing; if you are in a parking garage and you unlock with your fob and another car unlocks at the same time or you slam your car door and another car alarm goes off.)}}
# (5 min) How might using just one of Hill's criteria lead someone to a wrong conclusion? {{BoxAnswer|For each causal link you can work through the different criteria (i.e. Every time you press the unlock button on your car keys it only unlocks your car, this mechanism works at home, in a parking garage, and on a street, the mechanism works in any context such as rain/sun/night/day, other people can do it, etc.)}}
 
</div></div> -->
 
{{NavCard|prev=6.1 Correlation and Causation|next=7.1 Causation, Blame, and Policy}}
[[Category:Lesson plans]]
[[Category:Lesson plans]]

Latest revision as of 23:00, 11 June 2026

In the messy real world, an ideal randomized controlled trial may not always be feasible, for ethical or practical reasons. Even so, it is still possible to present compelling evidence for causation by considering whether the observed data satisfy a set of intuitive criteria introduced by Bradford Hill.

The Lesson in Context

This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.

Earlier Lessons

2.2 Systematic and Statistical Uncertainty
  • When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.
3.1 Probabilistic Reasoning
  • A credence level is always associated with any scientific claim of causation. As each Hill's criterion adds to a case for causation, so our credence level for a causal relation increases. However, unlike in the case of RCTs, it may be difficult to quantify.
6.1 Correlation and Causation
  • RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.

Later Lessons

11.2 When Is Science Suspect
  • Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.

Takeaways

After this lesson, students should

  1. Identify cases in which "ideal" RCT experiments are not possible, due to ethical or practical constraints.
  2. For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
  3. Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and consistency across contexts.
  4. Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.

Hill's criteria

A group of criteria that suggest possible causation even in the absence of an RCT.
  • Prior Plausibility
Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
  • Temporality/Temporal Sequence
Did the hypothesized cause precede the effect?
  • Specificity
Specific predictions for specific consequences that have come true are less likely to be caused by other factors.
  • Dose-response Curve
Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?
  • Consistency Across Contexts
Does the correlation appear across diverse contexts?

This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.

It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).


Leaded Gasoline and Violent Crimes

The causal connection between leaded gasoline across the world and violent crimes in many countries.
  • Prior Plausibility: High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
  • Temporality: In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
  • Specificity: Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
  • Dose-response Curve: See temporality.
  • Consistency Across Contexts: The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.

Exemplary Quotes

Ok, I agree that 'correlation doesn't prove causation' in general, but in a case like this where we have lots of other kinds of evidence it sure gives us a pretty strong guess about causation.

About a third of people admitted to relationship cycling, with similar rates among couples of different sexual orientations. That behavior, the researchers found, was correlated with increases in psychological distress, even after accounting for other factors that can influence mental health, such as demographic information, marital and family status, sexual orientation and related stressors. The more on-off cycles a person reported, Monk says, the larger the increases in depression and anxiety seemed to be.

From this, it follows that predicting future landscape patterns is difficult because contingencies may be unanticipated or even unpredictable. When similar locations can arise from different histories, and similar histories can produce different outcomes (e.g., Ernoult et al. 2006), it is not easy to infer causation.

Landscape Ecology in Theory and Practice, pg. 58

One hundred years after that, French chemist Antoine Lavoisier used a device called an "ice calorimeter" to gauge the energy burn from animals —like guinea pigs— in cages by watching how quickly ice or snow around the cages melted. This research suggested that the heat and gases respired by animals, including humans, related to the energy they burn.

I don't think there is a simple answer, because it is not clear that there is an increased risk of premature death - these are all observational studies and we are not controlling the amount of alcohol people consume and then analyzing the risk of death.

Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.

A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.

Students are quick to notice small sample size, slower to notice problems with experimental design.

Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.

After this lesson, students should

  1. Attitudes
    1. Appreciate that we can sometimes get very good evidence for a causal hypothesis, even in the absence of decisive RCTs.
    2. Be wary of potential confounds in apparent evidence for causality.
  2. Concept Acquisition
    1. In many cases it is not possible to conduct a true RCT to test causality, for practical or ethical reasons.
    2. There are non-RCT forms of evidence for causal hypotheses, which are less conclusive than RCTs but together can offer strong evidence for causation. These include:
      1. Prior plausibility: Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
      2. Temporality/temporal sequence: Did the hypothesized cause precede the effect?
      3. Dose-response curve: Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect?
      4. Consistency across contexts: Does the correlation appear across diverse contexts?
  3. Concept Application
    1. For a given causal hypothesis and imperfect study, identify the imperfections (e.g., sample size, lack of randomization, lack of control) and explain how these imperfections impact claims of causality.
    2. Identify potential confounds in RCT and non-RCT studies.
    3. Sketch out methods for eliminating potential confounds in sample RCT or non-RCT studies.
    4. For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
    5. For a given scenario in which a causal hypothesis is being made, describe an ideal experiment/set of experiments to test the hypothesis and rule out alternative hypotheses.
    6. Identify cases in which 'ideal' experiments are not possible, due to ethical or practical constraints.
    7. Evaluate the strength of causal claims when various sources of evidence are used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.
    8. Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, size of effect, temporal ordering, and multiple complementarily-flawed experiments.

Additional Content

You must be logged in to see this content.