11.2 When Is Science Suspect
More actions

The capacity for science to be misused to reinforce existing power structures.
The Lesson in Context
After discussing how science may go wrong in its form in 11.1 Pathological Science, we now discuss how science may go wrong in its societal outcome. Through a discussion activity, students will appreciate the difficulty in measuring humans in a way that is absolutely free of confounds that have discriminatory implications, and thus learn the importance of social responsibility as a scientist and including a more representative array of voices in selecting scientific questions and designing measures.
Takeaways
After this lesson, students should
- Recognize the potential to abuse science for social and political ends.
- Show heightened caution in situations in which science involves the study of human groups and subsequent validation of societal power structures.
- Recognize that you yourself are always involved in some social dynamic that may be relevant to the assessment of any particular study of human groups.
While it may be easy to spot cases where science has been used intentionally to validate preexisting societal power structures, the goal of this lesson is to bring attention to the possibility that one may inadvertently use science in such a way, through negligence or even with the best intentions, especially when it comes to the study of human groups.
Just World Fallacy
Reliability
Validity
External Validity
- Longevity of Effect
- Disruption Effect
- Novelty Effect
The Flynn Effect
- Scores from IQ tests are typically normalized such as to set the average score to be 100. But, throughout the 20th century, there was a consistent increase in how well people performed on IQ tests. This means that when newer test subjects take older tests, the average scores tended to be greater than 100.
Cognitive Capabilities by Ethnicity and Gender
- For a long time many scientists asserted that women and ethnic minorities were less intelligent than men of European descent, using studies of cognitive performance normalized to the skills such men spent their time developing. This claim was used to justify withholding educational and vocational opportunities for both women and minorities, as well as institutional and legal power.
Genetic Description of Pellagra
- In 1916, Charles Davenport asserted that the disease pellagra was genetic (racial). However, it is actually caused due to nutritional factors (a lack of vitamin B3).
The Hereditary Factor In Pellagra
The introduction from Davenport's book describing his racial theory of pellagra.
Representation and External Validity
- Historically, psychological or medical studies are often performed on predominantly white, wealthy, young males, with conclusions that are nevertheless generalized to the culturally and genetically diverse global population. For example, while blacks and Latinos make up 30% of the US population, they represent only 6% of participants in federally funded studies. While the studies themselves may not be discriminatory explicitly or implicitly, the differing validity of their application to diverse populations may be a source of furthered inequality.
Only a racist scientist could produce a racially discriminatory study.
Useful Resources
Recommended Outline
Before Class
Carefully review the case studies so that you're ready to answer any questions on them during the lesson.
During Class
| 10 Minutes | Go through the discussion questions. |
| 35 Minutes | Do the admissions officer activity. |
| 5 Minutes | Go through the happiness discussion. |
| 10 Minutes | Present the vocabulary tests example. |
| 10 Minutes | Present several case studies of inclusionary research agendas. |
| 10 Minutes | Run the closing discussion. |
Lesson Content
Discussion Questions
Break up into small groups for this activity.
- What groups are you a part of? Academic, social, cultural, interest-based groups? Are you a part of multiple groups? What groups do you have in common?
- What is your most salient (noticeable, important) group in the eyes of other people?
- What about in your own eyes?
- Why do we "talk up" and think better of groups we are in? What experiences have you had thinking better of a group (e.g. your high school sports team, home room, or grade cohort) because you were a part of it?
- Should we give up altogether on the study of human groups, or is there important work to be done if we are careful?
Note that we are not just talking about ethnic groups.
Admissions Officer
We want to demonstrate how difficult it is to come up with criteria with which to measure humans without inadvertently disadvantaging certain groups of people. Nevertheless, these imperfect criteria are often better able to select for certain traits than picking people at random. Having some form of criteria is often necessary. In this activity the students are split into small groups and have to act as admissions officers for various groups and teams. The tests are meant to capture the most important aspects of one's aptitude in a certain task, but due to limited time and resources, these tests are inevitably highly flawed.
Instructions
| 6 Minutes | Explain the activity. |
| 7 Minutes | The students discuss and write their criteria down. |
| 7 Minutes | The students present their criteria to a paired group and listen to the other group's criteria. |
| 7 Minutes | The students brainstorm why the other group's specific criteria may inadvertently exclude kids that are still really good "players" if judged holistically but wouldn't score well on those specific criteria. |
| 4 Minutes | Have your class go over the admissions discussion questions. |
| 4 Minutes | Have your class discuss the activity's relevance to science with the AOT questions. |
Explanation
Students will work in small groups. As an admissions officer, they need to select 20 out of 300 kids. You only have one day of "tryouts" or "test", for 5 minutes per applicant. You can give them any test/task you want, but at the end of the day, the choices must depend only on the result of this test. There can be many test items/questions/tasks, but each applicant needs to receive one numerical value at the end of their 5-min test.
Reminder
The 20 applicants with the highest scores are admitted. This test is the only criterion that will be used to judge the applicants.
Tell your students to do their best to select people who have the highest aptitude in the given prompt. Clearly write down the task and criteria for judging.
Prompts
- American football team
- Chess (note a full game typically takes 10 min)
- Children's choir
- Honours mathematics programme
- Programming camp
- Rock climbing
- Rowing
- E-sports league (e.g. Smash bros)
- Model UN
- History camp
- Poetry group
Admissions Discussion Questions
- What are the difficulties that you faced when coming up with selection criteria?
- Why might an otherwise highly qualified kid be rejected?
- How would you improve your criteria or the selection process itself to be more inclusive?
Reflection on AOT
"My blood boils over whenever a person stubbornly refuses to admit he's wrong." This question is from the AOT test from 3.2 Calibration of Credence Levels.
- What should we be mindful of when we do and interpret this study, especially for comparison between human groups?
- Should we not have given you the AOT test?
Takeaways
- Designing valid criteria to measure humans is difficult. Often, certain groups are disadvantaged in these measurements.
- Nevertheless, even imperfect criteria have advantages over instinct, nepotism, randomness... Some sort of criteria often seems necessary.
- We can accept that a test will be highly flawed, while still striving for a valid and equitable test. We can do our best on this impossible (but necessary) task.
Measuring Happiness
This discussion brings up a metric that people often do want to study in the real world and the cross-cultural difficulties inherent in doing so.
Initial Happiness Discussion Questions
Students should discuss the following questions in small groups.
- In your group, what are all the languages that you collectively speak or understand?
- What are all the words for "happiness" in these languages?
- Do they mean the same thing?
- What are the nuanced differences between the meanings of these different words for "happiness?"
- If you wanted to compare the levels of happiness among speakers of different languages, how might the different words affect the results of these studies?

Then show the students the table. It was put together by the Japanese researchers Uchida and Ogihara to emphasize the cross-cultural differences in meanings of happiness that make it challenging to compare across different countries. Afterwards, ask the students to discuss the following prompt.
- Suppose you are in charge of an international scientific funding institution and want to promote projects that measure and compare "happiness" across different languages, contexts, and cultures.
- You want the studies to be as reliable and valid as possible.
- What conditions would you put in place or look for in the research you fund to support this?
There isn't any one answer to this question. But, we're hoping the students come up with the idea of needing researchers from a diverse set of backgrounds from the very beginning of study design.
Activity 1
[Brief description of and motivation for the activity]
[Common misconception or thing to look out for.]
[Thing you really need to look out for!]
[Title]
[Useful tip, guideline, or other background.]
Instructions
| [n] Minutes | [Activity.] |
| [n] Minutes | [Activity.] |
Discussion Questions
[Question 1]
[Possible misconception that may need to be corrected and clarified.]
[Intended answer to the above question.]
[Question 2]
[Possible misconception that may need to be corrected and clarified.]
[Intended answer to the above question.]