11.2 When Is Science Suspect
More actions

The capacity for science to be misused to reinforce existing power structures.
The Lesson in Context
After discussing how science may go wrong in its form in 11.1 Pathological Science, we now discuss how science may go wrong in its societal outcome. Through a discussion activity, students will appreciate the difficulty in measuring humans in a way that is absolutely free of confounds that have discriminatory implications, and thus learn the importance of social responsibility as a scientist and including a more representative array of voices in selecting scientific questions and designing measures.
Takeaways
After this lesson, students should
- Recognize the potential to abuse science for social and political ends.
- Show heightened caution in situations in which science involves the study of human groups and subsequent validation of societal power structures.
- Recognize that you yourself are always involved in some social dynamic that may be relevant to the assessment of any particular study of human groups.
While it may be easy to spot cases where science has been used intentionally to validate preexisting societal power structures, the goal of this lesson is to bring attention to the possibility that one may inadvertently use science in such a way, through negligence or even with the best intentions, especially when it comes to the study of human groups.
Just World Fallacy
Reliability
Validity
External Validity
- Longevity of Effect
- Disruption Effect
- Novelty Effect
The Flynn Effect
- Scores from IQ tests are typically normalized such as to set the average score to be 100. But, throughout the 20th century, there was a consistent increase in how well people performed on IQ tests. This means that when newer test subjects take older tests, the average scores tended to be greater than 100.
Cognitive Capabilities by Ethnicity and Gender
- For a long time many scientists asserted that women and ethnic minorities were less intelligent than men of European descent, using studies of cognitive performance normalized to the skills such men spent their time developing. This claim was used to justify withholding educational and vocational opportunities for both women and minorities, as well as institutional and legal power.
Genetic Description of Pellagra
- In 1916, Charles Davenport asserted that the disease pellagra was genetic (racial). However, it is actually caused due to nutritional factors (a lack of vitamin B3).
The Hereditary Factor In Pellagra
The introduction from Davenport's book describing his racial theory of pellagra.
Representation and External Validity
- Historically, psychological or medical studies are often performed on predominantly white, wealthy, young males, with conclusions that are nevertheless generalized to the culturally and genetically diverse global population. For example, while blacks and Latinos make up 30% of the US population, they represent only 6% of participants in federally funded studies. While the studies themselves may not be discriminatory explicitly or implicitly, the differing validity of their application to diverse populations may be a source of furthered inequality.
Only a racist scientist could produce a racially discriminatory study.
Useful Resources
Recommended Outline
Before Class
Carefully review the case studies so that you're ready to answer any questions on them during the lesson.
During Class
| 10 Minutes | Go through the discussion questions. |
| 35 Minutes | Do the admissions officer activity. |
| 5 Minutes | Go through the happiness discussion. |
| 10 Minutes | Present the vocabulary tests example. |
| 10 Minutes | Present several case studies of inclusionary research agendas. |
| 10 Minutes | Run the closing discussion. |
Lesson Content
Discussion Questions
Break up into small groups for this activity.
- What groups are you a part of? Academic, social, cultural, interest-based groups? Are you a part of multiple groups? What groups do you have in common?
- What is your most salient (noticeable, important) group in the eyes of other people?
- What about in your own eyes?
- Why do we "talk up" and think better of groups we are in? What experiences have you had thinking better of a group (e.g. your high school sports team, home room, or grade cohort) because you were a part of it?
- Should we give up altogether on the study of human groups, or is there important work to be done if we are careful?
Note that we are not just talking about ethnic groups.
Admissions Officer
We want to demonstrate how difficult it is to come up with criteria with which to measure humans without inadvertently disadvantaging certain groups of people. Nevertheless, these imperfect criteria are often better able to select for certain traits than picking people at random. Having some form of criteria is often necessary. In this activity the students are split into small groups and have to act as admissions officers for various groups and teams. The tests are meant to capture the most important aspects of one's aptitude in a certain task, but due to limited time and resources, these tests are inevitably highly flawed.
Instructions
| 6 Minutes | Explain the activity. |
| 7 Minutes | The students discuss and write their criteria down. |
| 7 Minutes | The students present their criteria to a paired group and listen to the other group's criteria. |
| 7 Minutes | The students brainstorm why the other group's specific criteria may inadvertently exclude kids that are still really good "players" if judged holistically but wouldn't score well on those specific criteria. |
| 4 Minutes | Have your class go over the admissions discussion questions. |
| 4 Minutes | Have your class discuss the activity's relevance to science with the AOT questions. |
Explanation
Students will work in small groups. As an admissions officer, they need to select 20 out of 300 kids. You only have one day of "tryouts" or "test", for 5 minutes per applicant. You can give them any test/task you want, but at the end of the day, the choices must depend only on the result of this test. There can be many test items/questions/tasks, but each applicant needs to receive one numerical value at the end of their 5-min test.
Reminder
The 20 applicants with the highest scores are admitted. This test is the only criterion that will be used to judge the applicants.
Tell your students to do their best to select people who have the highest aptitude in the given prompt. Clearly write down the task and criteria for judging.
Prompts
- American football team
- Chess (note a full game typically takes 10 min)
- Children's choir
- Honours mathematics programme
- Programming camp
- Rock climbing
- Rowing
- E-sports league (e.g. Smash bros)
- Model UN
- History camp
- Poetry group
Admissions Discussion Questions
- What are the difficulties that you faced when coming up with selection criteria?
- Why might an otherwise highly qualified kid be rejected?
- How would you improve your criteria or the selection process itself to be more inclusive?
Reflection on AOT
"My blood boils over whenever a person stubbornly refuses to admit he's wrong." This question is from the AOT test from 3.2 Calibration of Credence Levels.
- What should we be mindful of when we do and interpret this study, especially for comparison between human groups?
- Should we not have given you the AOT test?
Takeaways
- Designing valid criteria to measure humans is difficult. Often, certain groups are disadvantaged in these measurements.
- Nevertheless, even imperfect criteria have advantages over instinct, nepotism, randomness... Some sort of criteria often seems necessary.
- We can accept that a test will be highly flawed, while still striving for a valid and equitable test. We can do our best on this impossible (but necessary) task.
Measuring Happiness
This discussion brings up a metric that people often do want to study in the real world and the cross-cultural difficulties inherent in doing so.
Initial Happiness Discussion Questions
Students should discuss the following questions in small groups.
- In your group, what are all the languages that you collectively speak or understand?
- What are all the words for "happiness" in these languages?
- Do they mean the same thing?
- What are the nuanced differences between the meanings of these different words for "happiness?"
- If you wanted to compare the levels of happiness among speakers of different languages, how might the different words affect the results of these studies?

Then show the students the table. It was put together by the Japanese researchers Uchida and Ogihara to emphasize the cross-cultural differences in meanings of happiness that make it challenging to compare across different countries. Afterwards, ask the students to discuss the following prompt.
- Suppose you are in charge of an international scientific funding institution and want to promote projects that measure and compare "happiness" across different languages, contexts, and cultures.
- You want the studies to be as reliable and valid as possible.
- What conditions would you put in place or look for in the research you fund to support this?
There isn't any one answer to this question. But, we're hoping the students come up with the idea of needing researchers from a diverse set of backgrounds from the very beginning of study design.
Vocabulary Tests

This is a presentation of preliminary research from a former SSS instructor that demonstrates how even metrics which seem culturally impartial might surprise us in how they don't generalize. Cross-cultural vocabulary tests like the Peabody Picture Vocabulary Test (PPVT) uses simple images and has young kids identify them when given prompts in their native language. The test is supposed to be able to be applied universally. But, it turns out that there are contexts like rural (and semi-rural) Kenya where kids may have limited exposure to picture books and other iconography.

Preliminary research suggests that these kids without as much exposure to images perform worse on picture-based vocabulary tests.

However, this doesn't mean that the kids actually have worse vocabularies. When given actual objects instead of images of them, the kids perform substantially better.
The objects were all selected to be ones that the toddlers would encounter in everyday life. The images were drawn to match the PPVT style but be otherwise identical to the 3D objects.
Details About Study 1
- 32 toddler participants (mean = 2.29 years, SD = 0.21 years, 2.01 - 2.87 years)
- 5 additional toddlers tested (4 fussouts, 1 experimenter error)
- Tests were performed by local experimenter that speaks Luo, the first languages of the kids.
- Study was done following best research practices for cross-cultural psychology.
- Local collaborators, community involvement, timely dissemination of results to participating schools/parents, etc.
Details About Study 2
- Preregistered sample of n = 192 (96 children per condition)
- All children in their first year of formal schooling
- Within sample, lots of diversity in picture experience
- M = 4.56 years, SD = .96 years, range = 2.38 - 7.48 years
- Gave the kids PPVT-style vocabulary assessment in Swahili, involving either pictures or objects
- For half the kids give them pictures, for the other half gave them objects
- Note that figure includes the combined results for nouns, numbers, and colors
- The kids do very well with nouns overall
- The kids have some trouble with numbers - this is common across cultural contexts
- The kids really have difficulty with colors. They rarely perform better than chance. This is because colors are not something normally taught in this part of Kenya. And, when it is taught, its only much later, in schools, and in English. Many of the (adult) local collaborators also had difficulty identifying colors in their otherwise native language.
- Note that this test was done in Mombasa.
- Mombasa is the second-largest city in Kenya and is a much more urban environment.
- That being said, this was done in one of the more rural parts of Mombasa.
- Almost all the schools in this area have murals on their walls. So, theres definitely some exposure to iconography.
Research Agendas
This section is a series of case studies of different researchers from historically underrepresented backgrounds. Their research direction was shaped by their background and, in all cases, wasn't a priority for other researchers in the field. We hope to address the following points.
- Having researchers from diverse backgrounds can affect the quality of cross-cultural science being done.
- How can it also affect which scientific questions are studied in the first place?
- What are the boundary conditions for generalization? If results replicate in two cultures? Every single culture?
- If most populations respond one way and there is one exception, what kind of conclusions can be drawn?
Case Study 1

This document was made for a different iteration of the course. Some of the discussion in it may not apply.
- The rivers near Africa Flores' village in Guatemala were all too polluted from sewage and agricultural chemicals to be potable.
- However, she discovered the water in other parts of the country was clean.
- Ultimately discovered that the discrepancy was due in part to a lack of information about water quality in different areas.
- Data on Guatemalan water quality was already being taken by the U.S. Geological Survey (USGS) and NASA since 1970.
- In 2008, that data was finally open sourced.
- Flores took ground samples to calibrate the satellite data and track changes in the 2009 Lake Atitlan toxic algal bloom.
- She works for NASA to track environmental changes in data-poor countries.
Case Study 2

- Native Alaskans have been historically excluded from academia and academic texts.
- Combines her Arctic experiences and Inuit perspective with theory and data to develop effective methods that equip others in planning and conducting respectful research, teaching, organizational strategy, and advocacy work.
- Assistant professor of professional and technical writing at Virginia Tech who focuses on empowerment, social justice, and equitable research practices.
Case Study 3
- Kaposi's Sarcoma-Associated Herpesvirus causes cancer in immunocompromised individuals (usu. HIV/AIDS patients).
- Originally characterized in part by researchers at UCSF after the HIV pandemic.
- Now primarily endemic in Africa where HIV positivity rates are high.
- Schism between what is studied in the US and other Western countries (basic biology) versus Africa-based groups pursuing epidemiology and clinically-focused research.
- Formerly, the annual KSHV conference was always in the US. Now to promote better collaboration and accessibility, the conference is in South Africa every other year.
Closing Discussion
What should scientists be mindful of when studying human groups or members of a different group?