4.2 Finding Patterns in Random Noise
More actions

We often find mistake noise for signal; how do we minimize these mistakes, given that they are not always easy to tell apart?
The Lesson in Context
This lesson continues 4.1 Signal and Noise by elaborating on ways in which random noise can emulate signals (produce apparent patterns) in many different contexts. We introduce the idea of [math]\displaystyle{ p }[/math]-values to quantify the statistical significance of patterns and describe various tempting statistical fallacies we tend to make as laypersons or scientists, such as gambler's fallacy and [math]\displaystyle{ p }[/math]-hacking. We play a game in which students try to produce a random string of coin tosses by thought, which reveals that a truly random string in fact contains more apparent patterns than one intuitively expects. Two other activities also illustrate how spurious patterns are in fact expected to arise from random noise.
Takeaways
After this lesson, students should
- Understand that people tend to see any regularity as a meaningful pattern (i.e., see more signal than there is), even when "patterns" occur by chance (i.e. are pure noise).
- Recognize cases of the Look Elsewhere Effect in daily life when you hear phrases such as "what are the odds".
- Recognize and explain the flaw in scenarios in which scientists and other people mistake noise for signal.
- Resist the opposing temptations of both the Gambler's Fallacy (the expectation that a run of similar events will soon break and quickly balance out, because of the assumption that small samples resemble large samples) and the Hot-hand Fallacy (the expectation that a run will continue, because runs suggest non-randomness).
People underestimate the frequency of apparent patterns produced by randomness, leading to over-perception of spurious signal much more frequently than people account for. Events that are just coincidental are much more likely than most people expect.
Look Elsewhere Effect
Since humans are so good at identifying noise that looks like signals, it is easy to find and fall victim to this even if it doesn't seem like we're asking too many questions. The look elsewhere effect can be avoided by clearly stating the questions you're asking before seeing the data.
Things that Cause the Look Elsewhere Effect
- Asking too many questions of the same data set, reporting only statistically significant results.
- Asking the same question of multiple data sets, reporting only statistically significant results.
- Running a test or similar tests too many times, reporting only statistically significant results.
- The effect also occurs in everyday life, e.g. when one looks at a whole lot of phenomena and only takes note of the most surprising-looking patterns, not properly taking into account the larger number of unsurprising patterns/lack of pattern.
[math]\displaystyle{ p }[/math]-hacking
A [math]\displaystyle{ p }[/math]-value cutoff of .05 thus indicates that, on average, 1 in 20 results will be false positives. So one should expect, on average, one false positive for every 20 independent analyses of pure noise. [math]\displaystyle{ p }[/math]-hacking is statistically problematic but more often a result of misunderstanding than deliberate fraudulence.
Common Techniques for [math]\displaystyle{ p }[/math]-hacking
- Running different statistical analyses on the same dataset and only reporting the statistically significant ones.
- Analyzing multiple DVs and only reporting the statistically significant ones.
- Gradually increasing the sample size until the [math]\displaystyle{ p }[/math]-value falls below .05.
[math]\displaystyle{ p }[/math]-hacking sounds malicious. But, it's easy to do inadvertently even as a professional researcher!
Gambler's Fallacy
Hot-hand Fallacy
These two fallacies lead in opposite directions, but are both a result of the misconception that small samples (or short runs) will resemble large samples (or long runs), forgetting about statistical uncertainty. The Gambler's Fallacy arises because people assume a small sample will look like a large sample, such that a run of e.g. Tails will quickly be balanced out by Heads. The Hot Hand Fallacy arises when a run (e.g. of Tails) makes people think the sequence isn't truly random, but an effect of skill or luck that will continue. Both fallacies arise because long runs are commoner in random sequences than people expect.
File-drawer Effect
Mr. Goxx
- A hamster that actively manages a cryptocurrency portfolio by running in his "intention wheel" to determine what he's buying/selling and then goes through either a "BUY" or "SELL" decision-tunnel to decide what he's doing with it. As of october 2021, his portfolio was up nearly 30% from june when he started trading crypto. His decisions are streamed live on Twitch.
Cold Reading in a Crowd
- A medium shouts out to a crowd somewhat common names, such as William and Butler, and common ailments such as, "heart problem" or "passed in their sleep". It is very likely that a couple people in the crowd can match a majority of these features to someone in their life, making the claims seem impressively accurate (seemingly small [math]\displaystyle{ p }[/math]-value), but this ignores all the other members of the audience who could not match any of the claims. This is an illustration of [math]\displaystyle{ p }[/math]-hacking.
[math]\displaystyle{ p }[/math]-hacking Through Incrementing Sample Size
- "We conducted the study on 1000 participants, and our [math]\displaystyle{ p }[/math]-value is just slightly above 0.05. Let's recruit another 20 participants to see if our [math]\displaystyle{ p }[/math]-value can dip under 0.05." If they keep increasing their sample size by 20 each time, the [math]\displaystyle{ p }[/math]-value will fluctuate just from chance alone, possibly dipping below 0.05 even if there isn't a real signal. This is equivalent to only selecting a subsample of the data that would confirm a hypothesis, omitting that an even larger sample would have rejected the same hypothesis.
"What are the odds!
- Every time you hear this, be suspicious. The odds are probably higher than you'd think, if you take into account all the similar events that didn't include anything surprising.
Running Into Someone
- You're on vacation in London and you run into an old friend that was also there on vacation. It may be unlikely that this specific friend came to this exact spot on vacation. But, you're bound to run into someone you've met before on some trip at some point, especially if you associate with people likely to visit similar places.
COVID Origin
- What are the odds that COVID originated in a market where there just so happened to be a nearby major virology research institute and a worker went home sick just before the outbreak? Viral outbreaks are likely to occur in densely populated areas in major cities. Major cities tend to have major research institutes/hospitals that might have some virology component, especially in locations where there are higher risks of new viruses emerging like high population density and/or wet markets. Anything in the city could be considered nearby. Workers go home sick all the time. All in all, it is insufficient to conclude whether or not COVID originated from that specific lab based solely on the perceived unlikeliness of the circumstances.
The Look Elsewhere Effect means that more data leads to more misinterpretations, or "too much data is bad."
More separate analyses someone runs, the better their analysis will be.
Wow, that was such a striking coincidence! It must have some hidden significance.
<restricted>
Useful Resources
Recommended Outline
Before Class
- Familiarize yourself with the Google Docs for the Fool the Professor game.
- Get enough pennies for the whole class for the Fool the Professor game.
- Familiarize yourself with the [math]\displaystyle{ p }[/math]-hacking notebook.
During Class
| 5 Minutes | Introduce the lesson and go over the plan for the day. Make sure people have groups, spokespeople, etc. |
| 24 Minutes | Go through the discussion questions. |
| 8 Minutes | Run the Stock Predictions activity. |
| 15 Minutes | Run the Fool the Professor activity. |
| 8 Minutes | Run the Snowy Pictures activity. |
| 20 Minutes | Use the remaining time to start work on the [math]\displaystyle{ p }[/math]-hacking notebook. |
After Class

One instructor should compile histograms for the Fool the Professor game.
- Separately for each of the two strings of coin tosses, tally the number of times a consecutive run of N heads or tails occurs. (This may be done with a script or the search-and-replace function in a text editor.)
- Mask which string is the true random one by labelling them A and B.
- Make the following two histograms for the tallies of consecutive runs, stacking the tallies for heads and tails. There's a calculator spreadsheet template to help you with this.
Lesson Content
Discussion Questions
Spend four minutes on each question.
Question 1

What's the joke in this comic?
If each test has a 5% chance of saying there's a signal when there's not, if you run enough tests (on enough colors of jelly beans), you're going to get one to say there's a signal even if there's only noise.
If your hypothesis is:
- "Do green jelly beans cause acne?" then you have a 5% chance of finding a false signal.
- "Do jelly beans cause acne?" then, with 20 experiments, you have a bout a 64% chance of finding a false signal. (1 - 0.95²⁰ ≈ 0.64).
Question 2
Why might humans be likely to see patterns in random noise?
Example Answers
Natural psychological discomfort with ambiguity and uncertainty, link back to superstitions in previous sections (seeing causation when there is only correlation). Discuss Patternicity article: when the costs of missing a real signal are higher than the costs of misperceiving a non-existent signal in noise, we are likely to perceive a signal. The example of seeing the "pattern" of a predator from the "noise" of a rustling patch of grass is a great example.
Question 3
What are some ways in which the Look Elsewhere Effect may impact scientific studies?
If a seemingly impressive "signal" appeared in one part of a large data set, a scientist might get very excited and publish that result, but they would've been equally impressed had the "signal" appeared elsewhere in the data set. The likelihood that some spurious signal appears anywhere in a large data set is quite high. The scientist has in this case forgotten that they should've looked elsewhere in the data set and noted that the "signal" in the totality of the data set is actually very weak.
Question 4
How do scientists try to stop themselves from seeing significant signals/patterns when only noise is present? (Hint: think [math]\displaystyle{ p }[/math]-values)
With tests of statistical significance! With a [math]\displaystyle{ p }[/math]-value cutoff of .05, you have a 5% chance of detecting a signal if there is truly only noise. You can adjust this threshold (typically .05) to suit your needs.
Question 5
What do you think of the results of the study in "Have smartphones destroyed a generation?" Do we trust it? Why or why not? How could we amend the analysis to trust it more?
It fell prey to the look elsewhere effect, since there were simply so many explanatory variables that is highly likely that one will be correlated purely by chance. We could correct this by having a lower [math]\displaystyle{ p }[/math]-value threshold, like .05/(number of variables considered)
Stock Prediction
Students guess whether each of four fictional stocks will rise or fall. The instructor picks if each stock will rise or fall by flipping a coin, and then asks the students if anyone got all four right. Typically, at least one student will, just by chance, even though it is clearly a matter of chance.
| Stock | Prediction | Results | Success |
|---|---|---|---|
| XMPL | Rise | Rise | Yes |
| KGYN | |||
| GFA | |||
| XGY | |||
| WIUT |
Instructions
| 1 Minute | Either hand out sheets of the table above or show it on a screen. For each of the stocks above, have the students predict whether they will "Rise" or "Fall" and write their predictions down. |
| 1 Minute | For each stock, flip a coin to determine if a stock will rise or fall. Do this in front of the students so they can see that the results really are random. |
| 1 Minute | Ask the students if anyone's predictions were all correct. With enough students, there's typically at least one student that gets everything right. |
| 3 Minutes | Praise any students that got everything correct. Emphasize how brilliant of traders they are and ask them what their methodology was. See if they have any profound insights about the market to share. |
| 2 Minutes | Remind students that the results here actually were truly random. If you look at a set of data in enough ways after the fact you'll inevitably find something that looks like a pattern. But this should demonstrate that the same thing applies if you make enough predictions ahead of time as well. |
Debrief
Did anyone do really well? That person must be an expert on stocks, right?
Nope. But, make sure you really ham it up and congratulate the students that do best.
Why did that person do really well?
Pure random chance. In a class of 30 students, we would expect one to three students to accurately predict the four coin flips by guessing randomly. There's a 1 in 16 chance of doing so.
What are some other cases in which luck looks like skill? What are consequences of this illusion?
Many possibilities. See the examples for this lesson.
Fool the Professor
This is a logistically heavy activity requiring some planning and an instructor (or volunteer) other than the professor.
This activity illustrates the fact that pure randomness can and will generate apparent patterns. Each class is split into two halves, with one half generating a random string of coin flips and the other half coming up with a string of coin flips that should "fool the professor" into thinking that that is the random string. The upshot is that the true random string tends to contain more long consecutive heads (or tails) than the fake random string.
Instructions
- The instructor should first familiarize themselves with the documents where the coin tosses will be recorded. Open the doc and be prepared to write down H or T. Prepare enough pennies to hand out to your students.
- Tell the students that the goal of this activity is to produce a string of coin tosses that can fool the instructor into believing that it is random, rather than produced by thought by students.
- Split the class into two halves (e.g. left and right halves of the room). Distribute the pennies to the second half of the class.
- The first half will produce a fake random string by shouting HEADS or TAILS one after another. After the GSI reads the last 10 coin tosses from the fake random string of the preceding section, let the students begin shouting their choices one by one. Record the results.
- The second half will produce a true random string by shouting the results of a real coin toss. After the GSI reads the last 10 coin tosses from the true random string of the preceding section, let the students shout the results of their coin tosses one by one. Record the results.
- The combined results of all sections will be revealed for the professor to guess. Typically, the true random sequence has longer strings of heads or tails than the fake-random sequence.
Manufactured Random Data Template
Document to put manufactured random data for the Fool the Professor activity.
You may allude to the ideas of false positives and false negatives, which will be introduced in the next lesson.
Extra Instructions for Multiple Simultaneous Sections

If the activity is done in multiple small sections, aggregate the two strings (one true random and one fake random) across all classes. The final two long strings are to be presented to the professor at the plenary session. To make sure the strings continue from those of the previous class, read the 10 trailing coin flips to the students before they generate the new strings.
Consider the list of coin flips listed from top to bottom. If there are two concurrent sections, then one section should add their coin flips to the bottom of the list, while the other section adds theirs to the top of the list (in the reverse order). See the image.
Snowy Pictures
In this activity, students will look at a series of "noisy" black and white images, each visible for only a second, with thirty second breaks between each for students to write down what they see. Some of the noisy pictures contain pictures of familiar creatures or objects, while others do not. The purpose of this activity is to give students practice looking for patterns where they may or may not exist, observing their responses as they "see" pictures that aren't there and fail to identify pictures that are really there because of too much noise.
You may allude to the ideas of false positives and false negatives, which will be introduced in the next lesson.
Instructions
| 8 Minutes | Show the students the pictures in these slides. For each image, have students look for a pattern or decide if it is only noise. Give ~5 seconds. After showing the picture, give thirty seconds for the students to write down what they saw or if it was just noise. |
| 4 Minutes | Once all pictures are shown, go through the pictures one more time. Before showing each image, ask the students to share what they saw. After getting answers from the students, pull up the picture and reveal whether or not it had anything in it. |
Answers
Which snowy pictures had images?
p-hacking Notebook
The last chunk of this lesson is reserved for students to work through the [math]\displaystyle{ p }[/math]-hacking Jupyter notebook in pairs. This is the first of several Jupyter notebooks we'll be using in this course. So, be ready to help the students in case they're confused. The students don't necessarily have to turn this in.
Many students without programming experience are put off by the code in the notebook. They don't need to understand any of it or know how to program at all. They only need to press "shift-enter" in any cells with code in them. However, spending some time to walk through it with them can also serve to be an empowering experience.
</restricted>