Jump to content
Toggle menu
Toggle preferences menu
Toggle personal menu
Not logged in
Your IP address will be publicly visible if you make any edits.

11.2 When Is Science Suspect

From Sense & Sensibility & Science
Revision as of 14:01, 14 August 2023 by Gpe (talk | contribs) (// Edit via Wikitext Extension for VSCode)

The capacity for science to be misused to reinforce existing power structures.



The Lesson in Context

After discussing how science may go wrong in its form in 11.1 Pathological Science, we now discuss how science may go wrong in its societal outcome. Through a discussion activity, students will appreciate the difficulty in measuring humans in a way that is absolutely free of confounds that have discriminatory implications, and thus learn the importance of social responsibility as a scientist and including a more representative array of voices in selecting scientific questions and designing measures.

1.1 Introduction and When Is Science Relevant
  • If science is to be used in societal decision making, scientists must be mindful of its potential to cause harm.
1.2 Shared Reality and Modeling
  • When studying humans, it is often necessary to define terms operationally, such as "what behaviors constitute altruism." Such operationalizations may be a crude measure of a more realist human trait, if one even exists. Cultural and personal biases may easily cause such definitions to favor the researchers' own group at the expense of others.
10.1 Confirmation Bias
  • It is tempting, though improper, for scientists to only publish results in a way that confirms the predominant belief. When it comes to human groups, this belief may be a mere stereotype.


Takeaways

After this lesson, students should

  1. Recognize the potential to abuse science for social and political ends.
  2. Show heightened caution in situations in which science involves the study of human groups and subsequent validation of societal power structures.
  3. Recognize that you yourself are always involved in some social dynamic that may be relevant to the assessment of any particular study of human groups.

While it may be easy to spot cases where science has been used intentionally to validate preexisting societal power structures, the goal of this lesson is to bring attention to the possibility that one may inadvertently use science in such a way, through negligence or even with the best intentions, especially when it comes to the study of human groups.


Just World Fallacy

The tendency to believe that outcomes are deserved and existing social structures are justified.

Reliability

The extent to which some metric is consistently measurable.

Validity

The extent to which some metric or scientific concept reflects some real external thing.

External Validity

The extent to which the result of an RCT generalizes to the world at large. There are several reasons results might not generalize.
  • Longevity of Effect
Lab studies tend to measure dependent variables immediately after the intervention, but often inferences are desired for long-term effects.
  • Disruption Effect
Awareness of being in an experiment.
  • Novelty Effect
Novelty of the context may alter the effects of the experiment.


The Flynn Effect

Scores from IQ tests are typically normalized such as to set the average score to be 100. But, throughout the 20th century, there was a consistent increase in how well people performed on IQ tests. This means that when newer test subjects take older tests, the average scores tended to be greater than 100.

Cognitive Capabilities by Ethnicity and Gender

For a long time many scientists asserted that women and ethnic minorities were less intelligent than men of European descent, using studies of cognitive performance normalized to the skills such men spent their time developing. This claim was used to justify withholding educational and vocational opportunities for both women and minorities, as well as institutional and legal power.

Genetic Description of Pellagra

In 1916, Charles Davenport asserted that the disease pellagra was genetic (racial). However, it is actually caused due to nutritional factors (a lack of vitamin B3).

Representation and External Validity

Historically, psychological or medical studies are often performed on predominantly white, wealthy, young males, with conclusions that are nevertheless generalized to the culturally and genetically diverse global population. For example, while blacks and Latinos make up 30% of the US population, they represent only 6% of participants in federally funded studies. While the studies themselves may not be discriminatory explicitly or implicitly, the differing validity of their application to diverse populations may be a source of furthered inequality.


Only a racist scientist could produce a racially discriminatory study.

As an example, a perfectly well-intentioned scientist studying the level of altruism in people across ethnic groups may come up with an operational definition of altruism that is based on their own cultural understanding of the term and fails to include alternative expressions of altruism in other cultures.

Useful Resources

Recommended Outline

Before Class

Carefully review the case studies so that you're ready to answer any questions on them during the lesson.

During Class

5 Minutes Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc.
10 Minutes Go through the discussion questions.
30 Minutes Do the admissions officer activity.
5 Minutes Go through the happiness discussion.
10 Minutes Present the vocabulary tests example.
10 Minutes Present several case studies of inclusionary research agendas.
10 Minutes Run the closing discussion.

Lesson Content

Discussion Questions

Break up into small groups for this activity.

  1. What groups are you a part of? Academic, social, cultural, interest-based groups? Are you a part of multiple groups? What groups do you have in common?
    • What is your most salient (noticeable, important) group in the eyes of other people?
    • What about in your own eyes?
  2. Why do we "talk up" and think better of groups we are in? What experiences have you had thinking better of a group (e.g. your high school sports team, home room, or grade cohort) because you were a part of it?
  3. Should we give up altogether on the study of human groups, or is there important work to be done if we are careful?

Note that we are not just talking about ethnic groups.

Admissions Officer

We want to demonstrate how difficult it is to come up with criteria with which to measure humans without inadvertently disadvantaging certain groups of people. Nevertheless, these imperfect criteria are often better able to select for certain traits than picking people at random. Having some form of criteria is often necessary. In this activity the students are split into small groups and have to act as admissions officers for various groups and teams. The tests are meant to capture the most important aspects of one's aptitude in a certain task, but due to limited time and resources, these tests are inevitably highly flawed.

Instructions

Students will work in small groups. As an admissions officer, they need to select 20 out of 300 kids. You only have one day of "tryouts" or "test", for 5 minutes per applicant. You can give them any test/task you want, but at the end of the day, the choices must depend only on the result of this test. There can be many test items/questions/tasks, but each applicant needs to receive one numerical value at the end of their 5-min test.

Reminder

The 20 applicants with the highest scores are admitted. This test is the only criterion that will be used to judge the applicants.


Do your best to select people who have the highest aptitude in the given prompt. Clearly write down the task and criteria for judging.

6 Minutes Explain the activity.
8 Minutes The students discuss and write their criteria down.
8 Minutes The students present their criteria to a paired group and listen to the other group's criteria.
8 Minutes The students brainstorm why the other group's specific criteria may inadvertently exclude kids that are still really good "players" if judged holistically but wouldn't score well on those specific criteria.

Prompts

  1. American football team
  2. Chess (note a full game typically takes 10 min)
  3. Children's choir
  4. Honours mathematics programme
  5. Programming camp
  6. Rock climbing
  7. Rowing
  8. E-sports league (e.g. Smash bros)
  9. Model UN
  10. History camp
  11. Poetry group

Discussion Questions

  1. What are the difficulties that you faced when coming up with selection criteria?
  2. Why might an otherwise highly qualified kid be rejected?
  3. How would you improve your criteria or the selection process itself to be more inclusive?

Reflection on AOT

"My blood boils over whenever a person stubbornly refuses to admit he's wrong." This question is from the AOT test from 3.2 Calibration of Credence Levels.

  1. What should we be mindful of when we do and interpret this study, especially for comparison between human groups?
  2. Should we not have given you the AOT test?

Takeaways

  • Designing valid criteria to measure humans is difficult. Often, certain groups are disadvantaged in these measurements.
  • Nevertheless, even imperfect criteria have advantages over instinct, nepotism, randomness... Some sort of criteria often seems necessary.
  • We can accept that a test will be highly flawed, while still striving for a valid and equitable test. We can do our best on this impossible (but necessary) task.

Activity 1

[Brief description of and motivation for the activity]

[Common misconception or thing to look out for.]

[Thing you really need to look out for!]

[Title]

[Useful tip, guideline, or other background.]


Instructions

[n] Minutes [Activity.]
[n] Minutes [Activity.]

Discussion Questions

[Question 1]

[Possible misconception that may need to be corrected and clarified.]

[Intended answer to the above question.]


[Question 2]

[Possible misconception that may need to be corrected and clarified.]

[Intended answer to the above question.]