Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

10.2 Blinding: Difference between revisions

From Sense & Sensibility & Science
// Edit via Wikitext Extension for VSCode
No edit summary
Line 1: Line 1:
[[File:Topic Cover - 10.2 Blinding.png|right|400px]]
[[File:Topic Cover - 10.2 Blinding.png|thumb]]


Blind analysis, the practice of deciding how we will analyze data before finding out if the analysis we have chosen supports our hypothesis, counteracts confirmation bias.
Blind analysis, the practice of deciding how we will analyze data before finding out if the analysis we have chosen supports our hypothesis, counteracts confirmation bias.

Revision as of 16:10, 21 July 2023

Blind analysis, the practice of deciding how we will analyze data before finding out if the analysis we have chosen supports our hypothesis, counteracts confirmation bias.



Useful Links

Readings and Assignments

Lecture Video

Learning Goals

After this lesson, students should

  1. Be able to explain why blind analysis might be needed, by explaining the errors that can arise in its absence.
  2. Recognize when blind analysis is being used and explain what function it serves. Identify situations and decisions in which blind analysis would be useful.
  3. Be able to evaluate techniques (e.g., registered replication, adversarial collaboration, peer review)
    1. for ability to address confirmation bias, and
    2. in comparison to blind analysis.
  4. Propose how to use blind analysis for simple studies.

Definitions

  • Blind Analysis
    Making all decisions regarding data analysis before the results of interest are unveiled, such that expectations about the results do not bias the analysis. Usually co-occurs with a commitment to publicize the results however they turn out.
  • Preregistration
    A research group publicly commits to a specific set of methods and analyses before they conduct their research.
  • Registered Replication
    One or more research groups commit to a specific set of methods and procedures to replicate earlier work to see if they get the same results (typically with the input of the original research team). Results are publicized regardless of outcome.
  • Adversarial Collaboration
    Scientists with opposing views agree to all the details of how data should be gathered and analyzed before any of the results are known.
  • Peer Review
    New results are evaluated by other experts in the same field to determine whether they are valid. This only reduces confirmation bias insofar as reviewers don't share the same biases.

Examples

  • Fermilab's muon g−2 experiment performed highly precise measurements of the magnetic dipole moment of muons to test the theoretical predictions of the currently accepted model of elementary particles. Blinding is done by injecting a secret code into all of the data that would undergo analysis, so that the scientists involved would not make specific choices in the analysis in a way that makes the final value agree with the theoretical prediction. The secret code was kept in a physical locker, the opening of which was highly publicized in the announcement event. Once the data was "unscrambled", the result shows that there is indeed a sizeable deviation of the measured value from the theoretical prediction. (Short video about this process)
  • One way in which p-hacking could occur is to choose or alter the analysis method after one has seen the results of that method to be undesirable. As an example, suppose a psychologist performs an experiment with 100 participants, sees that the results are at a statistical significance of p = 0.06, just shy of the p < 0.05 threshold for publication. They then decide to recruit another 100 participants to "improve their results", finally leading to p = 0.04, good enough for publication. This is a form of p-hacking, as p-values can dip below 0.05 as one slowly increases the sample size simply by random chance. To guard against this phenomenon, the sample size of a study is a required item in the preregistration process. See a demo of this type of p-hacking on this page.

Common Misconceptions

  • Students may confuse blind analysis with a double blind experiment. The latter is used primarily in treatment testing in conjunction with a placebo, such that the patient is prevented from knowing whether they received the real treatment or placebo, and the doctor is also prevented from knowing this fact in order not to inadvertently reveal this fact to the patient through subtle signs. The former type of blinding applies to the analysis process once the data has been collected. In the case of treatment testing, blind analysis may be employed whether or not double blinding is.

Context

This lesson offers solutions to potential pitfalls in scientific studies raised in 10.1 Confirmation Bias and 4.2 Finding Patterns in Random Noise. We use a stick measurement activity to illustrate the effects of confirmation bias and to motivate techniques to reduce its effect, especially blind analysis. These techniques are not universally employed in all fields of science today, and students pursuing a scientific career are encouraged to introduce these techniques in their own work.

Before

4.2 Finding Patterns in Random Noise
Blind analysis is one way to prevent some forms of p-hacking. For example, choices to be made about a study, such as the exact statement of the hypothesis, definitions of terms, and statistical techniques, can be preregistered or made "blinded" to the data. The effects on the final result due to these choices may be hidden from the researcher during analysis. These techniques prevent motivated reasoning during analysis decisions.
10.1 Confirmation Bias
Blind analysis helps prevent scientists from the temptation to make choices that make the result more likely to confirm the researcher's own prediction or to match currently accepted knowledge. Otherwise, non-confirming, surprising results may be incorrectly missed.

After

13.1 Denver Bullet Study
In group decision making, when factual evaluation and values evaluation are made by two different groups of people without knowledge of the other, it prevents evaluation motivated by the need to confirm a personal belief.

Recommended Outline

Before Class

  • Prepare a seating chart.
  • Do the preparation for the kabob kalimba activity.
  • Review PlayPosit and discussion questions and ask faculty, Gabriel, or Emlen any questions you have.
  • (Optional) Prepare a presentation.

During Class

  • (5 min) Come up with some fun way to assign the roles of spokesperson and notetaker (e.g. earliest birthday in the year, lives furthest from campus). Remind them of the responsibilities of these roles.
  • (1 min) Inform the students of the papers to read for next week's pathological science lesson.
  • (74 min) Spend the rest of the day on the tube measurement activity.

After Class

  • Collect answers from notetakers for the forum / plenary.

Lesson Content

Kabob Kalimba

This activity gives the students a chance to try and make some scientific measurement while falling into or avoiding several of the pitfalls that blind analysis could help with. In groups of three, the students will be cantilevering wooden skewers off the edge of a table and measuring the lengths at which they produce notes of different frequencies. The students' ultimate task is to find the ratio of the lengths of skewer stick (measured from the edge of the table to the end of the stick) for two notes that are an octave apart. The actual value is √2≃1.41. But, there is substantial priming to suggest to the students that the value they should expect is 1.5. If the students think of applying the methods of blind analysis, they should be able to avoid falling into this trap.

Preparation

  1. Print out at least ten copies of the Kabob Kalimba worksheet.
  2. Acquire at least 100 wooden kabob sticks.
  3. Make sure you have some method to number the groups the students work in.
  4. Have at least one sheet of paper for students to write down their final answers on.

Instructions

  1. (5 min) Introduce and setup the activity.
    1. Go through the slides briefly introducing the activity.
      You may have to demonstrate the activity as well.
    2. Split the class into groups of three.
    3. Assign each group a number.
    4. Give every group a copy of the worksheet and three.
    5. Remind the students to try and whisper for this activity so as to not mess up other people's data.
  2. (30 min) Let the students work on the activity and try to get their measurements.
  3. (5 min) Debrief the students on the gimmick of the activity.
    1. Have every group write their final ratio answer down on your piece of paper along with their group number.
    2. Reveal the final answer and the "trick" of the activity.
      The "trick" is that there's lots of priming for the answer to be 1.5 (or 3/2) but the real value is about 1.41. The priming showed up in lots of ways. Most obviously, the chart on the spreadsheet is centered on 1.5 and worksheet suggests finding the ratio between the 2nd and 3rd notes.
    3. Display and go through the answers the students came up with and see if anyone used or could have benefitted from blind analysis.
  4. (40 min) Go through the discussion questions with the whole class.

Discussion Questions

(20 minutes)

Answer these questions as a whole class.

  1. (1 min) What do you think the correct answer is? (Hint: It is not 2.)
  2. (3 min) Which senses or measurement instruments did you use?
  3. (4 min) What is the signal you are trying to measure, and what are the sources of noise?
  4. (4 min) What are the sources of systematic and statistical uncertainty? How did you address them?
  5. (4 min) What are some choices you had to make in your measurement and analysis? What are some places into which confirmation bias could have crept?
  6. (4 min) Did you use blind analysis? If so, how did you do it?

(8 min) Then have the groups discuss the following in small groups:

  1. Is there anything you would do differently to guard against bias in the measurement process?
    There's lots of ways to do this. But, they're all based around division of labor. The person listening for the octaves should probably not be the person using a ruler and making the measurements. You also could have someone independently make rules with different new units that convert inches or centimeters in ways that are unknown to the measurer.
  2. Is there anything you would do differently to guard against bias in the analysis process?
    This also depends on division of labor. What's most important is that the person doing the analysis has no sense of that the actual ratio values are when they're deciding what data points to include or not. The person measuring may also add a secret number to their measurements that's not revealed to the analyzer until after the analysis work is complete.

(8 min) Have the students come back and share over what they came up with as a class.

Overflow

Extra content that's not currently part of the official lesson plan.

Tube Measurement

For an in-person version of this activity, students would blow some plastic tubes to produce pitches. A short (fixed-length) tube, a long (adjustable-length) tube, and some measurement tools are provided. They are asked to find the ratio of lengths of the long tube to the short tube when the pitches they produce are one octave apart. Since the tuning and measurement are both difficult, some variation is inevitable in the final result. When asked to report a final ratio, the students may be inclined to ignore certain outlying points in order to confirm the expectation that the ratio is equal to 2, even when the true ratio is closer to 2.1. (Refer to old handouts before 2019 for this version.) This lets students experience first-hand the pull of confirmation bias and motivates certain blind analysis techniques.

In the Zoom version of this activity, the tuning and measurement are done by the course staff and video recorded. Half the groups are given "unblinded" worksheets and the other half "blinded" worksheets. The unblinded worksheets contain the actual measured ratios and try to prime students to confirm the ratio of 2. The blinded worksheets contain the same measured ratios, but with a secret number X added to all the values, so that students have to make choices about which data points to include without seeing whether the result will be close to 2.

Instructions

  1. Send the worksheets to students. Half the groups should receive the blinded worksheet and the other half the unblinded worksheets. Do not refer to them as blinded/unblinded or tell them what the difference between them is.
  2. Send the video link above to students. They may refer to parts of the video during their data analysis, but they do not need to watch the whole video.
  3. (X min) Let students work on the task and offer assistance wherever needed.
  4. After X min, remind each group to agree on a single ratio.
  5. Ask each group to tell you their final ratio privately, without revealing to other groups. For the blinded groups, you need to tell them that the secret number X = 1.37.
    You may alternatively use this Google Form, filled by one person per group.
    Update Google Form link.
  6. Close all breakout rooms. Reveal the difference between the blinded/unblinded worksheets and share the results of the whole class. We expect that the blinded groups would report something close to 1.41, while the unblinded groups would report something close to 1.5 due to confirmation bias.
    It's okay if this result doesn't occur. You could say that SSS students have often defied common human psychology.
  7. Ask students in the unblinded groups to describe what they were thinking when they were doing the analysis. Then ask the blinded groups.

Discussion Questions

  1. What was the purpose of adding a secret number X to all the measured ratios for the blinded groups?
    It shifts all the numbers away from the expected value of 2, in order to prevent confirmation bias.
  2. Are there other places where confirmation bias may creep in? Can you come up with other methods (during measurement or analysis) to further reduce confirmation bias?