4.2 Finding Patterns in Random Noise
More actions
Humans are so good at identifying patterns that we often see them even when it is really noise in masquerade. When we think we have seen a pattern, how do we quantify the level of confidence correctly? We describe common pitfalls that lead to an overconfidence in an apparent pattern, some that even prey on the inattentive scientist!
The Lesson in Context
This lesson continues 4.1 Signal and Noise by elaborating on ways in which random noise can emulate signals (produce apparent patterns) in many different contexts. We introduce the idea of [math]\displaystyle{ p }[/math]-values to quantify the statistical significance of patterns and describe various tempting statistical fallacies we tend to make as laypersons or scientists, such as gambler's fallacy and [math]\displaystyle{ p }[/math]-hacking. We play a game in which students try to produce a random string of coin tosses by thought, which reveals that a truly random string in fact contains more apparent patterns than one intuitively expects. Two other activities also illustrate how spurious patterns are in fact expected to arise from random noise.
Takeaways
After this lesson, students should
- Understand that people tend to see any regularity as a meaningful pattern (i.e., see more signal than there is), even when "patterns" occur by chance (i.e. are pure noise).
- Recognize cases of the Look Elsewhere Effect in daily life when you hear phrases such as "what are the odds".
- Recognize and explain the flaw in scenarios in which scientists and other people mistake noise for signal.
- Resist the opposing temptations of both the Gambler's Fallacy (the expectation that a run of similar events will soon break and quickly balance out, because of the assumption that small samples resemble large samples) and the Hot-hand Fallacy (the expectation that a run will continue, because runs suggest non-randomness).
- (Data Science) Describe the difference between the effect size (strength of pattern) and credence level (probability that the pattern is real), and identify the role each plays in decision making.
People underestimate the frequency of apparent patterns produced by randomness, leading to over-perception of spurious signal much more frequently than people account for. Events that are just coincidental are much more likely than most people expect.
Look Elsewhere Effect
Since humans are so good at identifying noise that looks like signals, it is easy to find and fall victim to this even if it doesn't seem like we're asking too many questions. The look elsewhere effect can be avoided by clearly stating the questions you're asking before seeing the data.
Things that Cause the Look Elsewhere Effect
- Asking too many questions of the same data set, reporting only statistically significant results.
- Asking the same question of multiple data sets, reporting only statistically significant results.
- Running a test or similar tests too many times, reporting only statistically significant results.
- The effect also occurs in everyday life, e.g. when one looks at a whole lot of phenomena and only takes note of the most surprising-looking patterns, not properly taking into account the larger number of unsurprising patterns/lack of pattern.
[math]\displaystyle{ p }[/math]-hacking
A [math]\displaystyle{ p }[/math]-value cutoff of .05 thus indicates that, 1 out of 20 analyses of pure noise would discover a spurious signal. [math]\displaystyle{ p }[/math]-hacking is statistically problematic but more often a result of misunderstanding than deliberate fraudulence.
Common Techniques for [math]\displaystyle{ p }[/math]-hacking
- Running different statistical analyses on the same dataset and only reporting the statistically significant ones.
- Analyzing multiple DVs and only reporting the statistically significant ones.
- Gradually increasing the sample size until the [math]\displaystyle{ p }[/math]-value falls below .05.
[math]\displaystyle{ p }[/math]-hacking sounds malicious. But, it's easy to do inadvertently even as a professional researcher!
HARKing
Gambler's Fallacy
Hot-hand Fallacy
These two fallacies lead in opposite directions, but are both a result of the misconception that small samples (or short runs) will resemble large samples (or long runs), forgetting about statistical uncertainty. The Gambler's Fallacy arises because people assume a small sample will look like a large sample, such that a run of e.g. Tails will quickly be balanced out by Heads. The Hot Hand Fallacy arises when a run (e.g. of Tails) makes people think the sequence isn't truly random, but an effect of skill or luck that will continue. Both fallacies arise because long runs are commoner in random sequences than people expect.
File-drawer Effect
Mr. Goxx
Cold Reading in a Crowd
[math]\displaystyle{ p }[/math]-hacking Through Incrementing Sample Size
"What are the odds!
Running Into Someone
COVID Origin
Exemplary Quotes
“I know it seems super meaningful that we ran into each other in Australia, when neither of us live in Australia, but I guess the chances of running into someone you know at some point, if you travel a lot and know a lot of people, are pretty high.”
“There are many many cases of people making insanely correct predictions, so many that some people are convinced clairvoyance is real. But there are many more cases of people making totally wrong predictions. So it's probably just noise; with enough predictions, someone will be correct by luck.”
The Look Elsewhere Effect means that more data leads to more misinterpretations, or "too much data is bad."
More separate analyses someone runs, the better their analysis will be.
Wow, that was such a striking coincidence! It must have some hidden significance.
Additional Content
You must be logged in to see this content.
