Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

6.2 Hill's Criteria

From Sense & Sensibility & Science
Revision as of 11:41, 15 August 2023 by Gpe (talk | contribs) (// Edit via Wikitext Extension for VSCode)

Building on Correlation and Causation, we examine how to collect evidence for causality in more difficult cases.



The Lesson in Context

This is a discussion-based lesson that familiarizes students with the concept of Hill's criteria, which are used when an ideal experiment (e.g. RCT) could not be done due to resource or ethical considerations. The criteria themselves are not difficult, but students typically have trouble associating their names with their meanings, and they would benefit from a diverse range of illustrative examples.

2.2 Systematic and Statistical Uncertainty
  • When coming up with alternative explanations to an apparent correlation between two variables, it helps to consider factors that contribute to systematic and statistical uncertainties.
3.1 Probabilistic Reasoning
  • A credence level is always associated with any scientific claim of causation. As each Hill's criterion adds to a case for causation, so our credence level for a causal relation increases. However, unlike in the case of RCTs, it may be difficult to quantify.
6.1 Correlation and Causation
  • RCTs are introduced as ideal experiments for establishing causal relationships. These are often not possible due to resource or ethical considerations, and we must resort to Hill's criteria to examine the plausibility of causation. Causation is defined as correlation under intervention. When manual intervention is not possible, Hill's criteria can still build a strong case for causation.
11.2 When Is Science Suspect
  • Students will explore how science has sometimes been used to justify the oppression of certain human groups. For example, current differences in achievement between human subgroups have been used to infer fundamental differences in biological or cognitive capacity, when in fact they may be sufficiently explained by differences in opportunity. These socially problematic causal inferences stem in part from the impossibility of an RCT that intervenes on genetics while keeping social opportunities equal between subgroups. It is important to recognise potential social implications when evaluating non-RCT evidence for causation, such as when drawing conclusions from observational studies alone.


Takeaways

After this lesson, students should

  1. Identify cases in which "ideal" RCT experiments are not possible, due to ethical or practical constraints.
  2. For a given scenario in which a causal hypothesis/claim is being made, identify plausible alternative hypotheses that could be consistent with the data.
  3. Identify additional sources of evidence that could be used to help mitigate flawed experiments, including prior plausibility, dose-response relationships, specificity, temporal ordering, and consistency across contexts.
  4. Recognize when causal evidence in the absence of an RCT can be fairly compelling, especially if there are many different types of evidence combined.

Hill's criteria

A group of criteria that suggest possible causation even in the absence of an RCT.
  • Prior Plausibility
Can a plausible mechanism be constructed, or is there some other basis for interpreting the current evidence in terms of one causal structure over another, such as data from other studies?
  • Temporality/Temporal Sequence
Did the hypothesized cause precede the effect?
  • Specificity
Specific predictions for specific consequences that have come true are less likely to be caused by other factors.
  • Dose-response Curve
Do the quantities of the hypothesized cause correlate with the quantity, severity, or frequency of the hypothesized effect across ranges?
  • Consistency Across Contexts
Does the correlation appear across diverse contexts?

This list may differ from the one on Wikipedia or elsewhere. It may be worth mentioning to students that these are the criteria we have chosen to focus on in this course.

It is not necessary for all of Hill's criteria to be satisfied to infer causation. Each criterion adds to the case for causation. Some criteria are not applicable in certain situations (e.g. dose-response curve in whether light switches cause the light to turn on and off).


Leaded Gasoline and Violent Crimes

The causal connection between leaded gasoline across the world and violent crimes in many countries.
  • Prior Plausibility: High levels of lead are known to cause cognitive damage. It is conceivable that extended exposure to lower levels of lead could have similar effects.
  • Temporality: In the graph shown in the video, the levels of violent crime correlate with the use of leaded gasoline, but delayed by about 20 years.
  • Specificity: Within the same demographic group, delinquents are 4 times more likely to have elevated bone lead concentrations than non-delinquents.
  • Dose-response Curve: See temporality.
  • Consistency Across Contexts: The delayed rise in violent crime after increased use of leaded gasoline is observed in many industrialized countries.

Since we can't run a randomized controlled trial on whether CO2 emissions cause global warming, we can't ever know whether it does.

A non-RCT study can still be serve as potentially weaker, but sometimes just as strong, evidence for causation.

Students are quick to notice small sample size, slower to notice problems with experimental design.

Students struggle to generate non-RCT types of evidence for causality, although they are better at recognizing it.

Useful Resources




During Class

5 Minutes Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc.
15 Minutes Ask your students the concept review questions.
25 Minutes Have your students work through the second part of the paper analysis activity in small groups.
15 Minutes Have your students discuss how to apply Hill's criteria to the social jetlag study.
15 Minutes Go over the COVID-19 example.
5 Minutes Answer any lingering questions. These activities have a tendency to go long, so some extra wiggle room is left at the end.

Lesson Content

Concept Review

What are the three essential elements of a Randomized Controlled Trial?

Random assignment, control group/condition, trial/experimental intervention

Note that random sampling is not essential to the validity of an RCT. It merely affects the generalizability of the results of the RCT.

For both of the following hypotheses, what RCT would you use to test it? Why would it be challenging to perform?

  1. For a child, does spending more than 22 hours in their home per day over their childhood increase their risk of developing myopia?

It is not hard to imagine an ideal RCT for this, but it is unethical/impractical to perform the intervention.

  1. On the level of individual towns, does implementing universal basic income reduce the rate of violent crimes in 10 years? In 150 years?

It is conceivable to do an RCT over 10 years, albeit expensive. However, it is hard to do this experiment on a large enough sample. It is also much less practical to track such an experiment over 150 years. It is difficult to randomise the towns.

Paper Analysis (Part 2)

This is a continuation of last lesson's paper analysis, in which we looked at a classic example of a proper RCT. Here, we will look at two more papers in which an RCT is difficult to conduct, but a strong case can still be made about causation based on experimental results.

Strictly speaking, the experiments in the following two papers are not RCTs, since the experimenter never performed any intervention. However, their causal conclusions are nearly as strong as an RCT. Hill's criteria, on the other hand, are useful when even this kind of experiment is impossible to do. We will go through an example of Hill's criteria in the COVID-19 example below.

For each of the two papers, students should try to answer the same questions as they did for the paper in part 1.

Vietnam War Draft Lottery

Example of natural experiment, where the intervention and assignment are not performed by the experimenter but by environmental factors.

The students should try to answer the following.

  1. What causal relationship is this paper trying to study? What is the hypothesis?

Does being drafted to serve in the military cause one's income to change later in life?

  1. What, if it exists at all, is the experimental intervention? (i.e., What is the independent variable that is being manipulated)?

Whether one has served in the military (veteran status).

  1. What is the dependent variable that is being measured (i.e., the variable the researchers anticipate may be affected by the experimental intervention)?

One's lifetime income.

  1. How is the independent variable manipulated? Are there control and intervention groups? Are those randomly assigned? Is the control condition a good one (only the independent variable is different, with all else kept equal)? Is this an RCT?

The veteran status is manipulated by the Vietnam War draft lottery, which randomly selected young men to serve in the military. The control group consists of young men who were not drafted by the lottery. This seems like a good control, as the selection is supposedly not based on any factor that may introduce systematic bias to one's future income. There are other factors that allow people to get out of the draft. This paper attempts to account for this. This is a natural experiment.

  1. What is the result of the experiment?

Veterans earn 15% less than non-veterans long after the draft and the war ended.

  1. Can the causal relationship in question 1 be concluded from the experimental results? If not, what, if anything, can be concluded? How confident are you in this conclusion?

Being drafted to serve in the military causes a long term reduction in one's income.

  1. Can you think of an alternative explanation for the data?

Well-connected or wealthier young men may have had an easier time getting out of the draft, which is a possible confound. Unknown confounds are more likely in a natural experiment, where the differences between experimental group and control group are not as tightly controlled as in a deliberately created RCT.

Upshot

This is a natural experiment. All the aspects of an RCT are still satisfied, but the randomized selection and intervention process is not performed by the experimenter but by an external agent/natural process.

Sunset Time and Social Jetlag

This paper is not technically an RCT but could be argued to have comparable strength.

For each of the two papers, students should try to answer the following.

  1. What causal relationship is this paper trying to study? What is the hypothesis?

Does an extra hour of natural light in the evening (according to local time) cause a change in one's sleep duration?

  1. What, if it exists at all, is the experimental intervention? (i.e., What is the independent variable that is being manipulated)?

The time of sunset in one's local time zone.

  1. What is the dependent variable that is being measured (i.e., the variable the researchers anticipate may be affected by the experimental intervention)?

Sleep duration.

  1. How is the independent variable manipulated? Are there control and intervention groups? Are those randomly assigned? Is the control condition a good one (only the independent variable is different, with all else kept equal)? Is this an RCT?

People who live on either side of a time zone boundary are considered, since they live geographically close to each other but have a 1-hour difference in sunset time. We may treat those on the east side of a boundary to be the intervention group and those on the west side to be the control group. The people are certainly not randomly assigned to one group or the other, but it is presumed that there are no systematic differences in the populations on either side that would cause a big difference in sleep patterns, other than the time zones. This is not an RCT, but its strength in implying causal relationship may be argued to be comparable to that of an RCT.

  1. What is the result of the experiment?

Those living in a time zone with one extra hour of daylight in the evening (an earlier sunset time) sleep 19 minutes less on average.

  1. Can the causal relationship in question 1 be concluded from the experimental results? If not, what, if anything, can be concluded? How confident are you in this conclusion?

Living in a time zone with one extra hour of daylight in the evening causes a 19-minute reduction in average sleep duration.

  1. Can you think of an alternative explanation for the data?

This is a pretty convincing study, what with the huge representative Census sample and the arbitrary cutoff. However, it's possible that some existing geographic power/wealth differential enabled people one one side of the time zone split to set it at their advantage (e.g. people in cities), leaving the less powerful/wealthy on the darker side (e.g. people in more rural areas). If this occurred, the difference in sleep and health could be due to the pre-existing power or wealth differential, not the time zone.

Discussion Questions

In your small group, explain Hill's criteria using aspects of the Sunset Time and Social Jetlag study as examples.

Prior Plausibility

The presence of sunlight and one's daily schedule are well known to affect sleep patterns. If the work schedule is shifted by one hour due to time zone, it should affect sleep patterns as well.

Temporality/Temporal Sequence

People have lived on one side of the time zone border before they developed this sleep pattern.

Specificity

The observed effect in sleep times is drastic across the time zone borders, and this effect is not observed elsewhere. This effect was also only observed in people that had to get up early for work.

Dose-response Curve

The further away one lives from a time zone border (more accurate their sunset/work time is to the natural circadian rhythm), the less pronounced the effect on sleep times becomes. The effect is also absent among unemployed people. See Figure 6 in the paper.

Consistency Across Contexts

The effect on sleep times is observed across many different time zone borders scattered throughout the US.

COVID-19

Bad in the olden days, there was a lot of controversy regarding mask wearing policies to combat the spread of COVID-19. Suppose it were very difficult to conduct a randomized controlled trial to study the effect of community mask wearing on COVID spread.

Have your students work in small groups. For each of the Hill's criteria, what evidence, if observed, would support that criterion?

Prior Plausibility

Masks physically blocks one's mouth from emitting/receiving virus-ridden aerosols.

Temporality/Temporal Sequence

Do COVID cases reduce after a mask mandate has been instituted in an area/business/school?

Specificity

Is it the case that counties/businesses/schools that implemented a mask mandate have a dramatically reduced case rate compared to those that did not?

Dose-response Curve

Does the extent of decrease in COVID cases correlate with the percentage of mask wearers in an area?

Consistency Across Contexts

Is this reduction in cases after a mask mandate observed in many countries/areas/businesses/schools?