011 min
Correlation and Causation, at a Glance
- What it is: A number that says two things moved together β never why they did.
- Origin: Pearson formalized the coefficient in 1896; Hill added causal criteria in 1965.
- Confused with: A confounding variable β a hidden third cause behind both halves of a correlation.
- Watch for: A correlation can pass a p < 0.05 significance test and still be pure coincidence.
021 min
Why Correlation and Causation Matters
A product analytics dashboard can show that users who enable a new setting retain at twice the rate of everyone else. A growth team that reads this as proof the setting drives retention will ship it to every new user and spend engineering time scaling it β and still lose the retention gain, if the real cause was that only power users, who already retained well, were the ones who found the setting in the first place. The dashboard measured a real correlation. It said nothing about what happens to a random new user pushed into that setting.
The same confusion travels outside software. A founder reading that startups funded by a top-tier investor grow faster might credit the funding, when both the funding and the growth trace back to the same underlying strength of the founding team. Mistaking correlation for causation does not just produce a wrong sentence in a report β it routes a real budget, a real roadmap, or a real hire toward the wrong lever.
032 min
What Correlation and Causation Actually Measures
Correlation and causation describe two different relationships between two variables, and the entire concept is the gap between them. Correlation is a measured quantity: the Pearson correlation coefficient, written r, scores how closely two variables move together on a scale from -1 to 1. An r of 0 means the two show no linear relationship; an r of 1 means that every time one rises, the other rises by a proportional amount; an r of -1 means one falls exactly as the other rises. Squaring r gives rΒ², the share of one variable's year-to-year movement that lines up with the other's β a purely descriptive number about how closely two series track each other over whatever period they were both measured.
Causation is a claim about mechanism, not motion: that changing one variable would itself produce the change in the other, holding everything else constant. Establishing it normally takes an intervention β an experiment, or an A/B test, where the researcher changes one variable directly and watches whether the second one moves, against a comparison group that was left unchanged. A correlation, however large, is compatible with three different underlying pictures: X actually causes Y; Y actually causes X, which is reverse causation; or a third variable, a confounding variable, produces both X and Y independently, so they move together without touching each other at all. The correlation coefficient cannot, on its own, tell you which of the three is true.
041 min
Where the Term Comes From
The correlation coefficient β the number "correlation" refers to β was formalized by the statistician Karl Pearson in 1896, in a paper on regression and heredity that built on earlier, cruder measures from Francis Galton. Pearson gave the relationship a single number between -1 and 1, still called Pearson's r today.
The warning half of the concept has no single inventor, but its most influential statement of how to move past it came from the epidemiologist Austin Bradford Hill. In a 1965 address to the Royal Society of Medicine, Hill proposed nine viewpoints β among them strength, consistency, and temporality β that a researcher should weigh before treating an association as causal, describing them as "nine different viewpoints from all of which we should study association before we cry causation." He was explicit that meeting all nine was never a requirement, only a discipline for the judgment call that correlation alone cannot make.
052 min
Worked Example
tylervigen.com's Spurious Correlations project publishes real, sourced yearly datasets already run through Pearson's formula, which makes it possible to show the full arithmetic: real inputs, a real output, and a result a reader can check.
One entry tracks two count variables from 2001 through 2020. The first is the number of movies Nicolas Cage appeared in each year, sourced from The Movie DB. The second is the number of magnitude 7.0-7.9 earthquakes recorded in the US each year, sourced from the USGS. Run through the standard Pearson formula, the 20 paired yearly counts produce the numbers below.
| Statistic | Value | What it means |
|---|---|---|
| r (Pearson correlation coefficient) | 0.5025504 | A moderate positive relationship between the two counts, over 2001-2020 |
| rΒ² (coefficient of determination) | 0.2525569 | 25.3% of one series' year-to-year movement lines up with the other's |
| p-value | 0.024 | Below the usual 0.05 threshold β this result reads as "statistically significant" |
| n | 20 paired years | 2001 through 2020, one pair of counts per year |
The site's own explanation of the rΒ² figure states it plainly: this means 25.3% of the change in the one variable (i.e., Magnitude 7.0-7.9 Earthquakes in the US) is predictable based on the change in the other (i.e., The number of movies Nicolas Cage appeared in) over the 20 years from 2001 through 2020. That sentence is true as a description of the two number series. It is not evidence that Nicolas Cage's film schedule shifts tectonic plates, or that earthquakes drive a casting director's decisions. Nothing physically connects the two variables β the arithmetic found a real, checkable pattern in two series that happened to move together for two decades. A p-value below 0.05 rules out one specific kind of bad luck: that this exact pattern would turn up by chance if you tested this one pair of variables a single time. The project's own creator is blunt about why that is not enough here: he ran the numbers on hundreds of millions of variable pairs and published the ones that matched, which is a far larger source of luck than any single p-value can rule out.
061 min
How Correlation and Causation Shows Up in Product, Design, and Engineering Work
A growth team reviewing a cohort dashboard sees that users who finish onboarding within their first session retain three times better than users who do not, and reads that as proof the fast-onboarding flow should be pushed on every new user. The dashboard reports a real correlation between two things that already happened together in the same users: finishing onboarding fast, and staying. It says nothing about what would happen to a user who does not naturally rush through onboarding but is pushed to. That gap is exactly why product teams run an A/B test instead of reading the correlation off a cohort report β a test assigns the change and watches what happens, rather than observing who chose it.
Machine learning engineers meet the same gap under a different name. A model trained to predict which users will churn can post a suspiciously strong accuracy score because one of its input features is itself a downstream effect of churn rather than a cause of it β a form of data leakage that produces a correlation the model cannot use once it is deployed on users whose future has not happened yet. A model that fits every wiggle in its training data, real pattern and coincidence alike, is overfitting: some of what it learned was never more than the training set's own version of the earthquake example above.
071 min
When to Trust a Correlation β and When Not To
A correlation is strong enough to act on when the two variables were randomly assigned to different groups β an A/B test, a randomized trial, a randomized policy rollout β because random assignment rules out confounding and reverse causation by construction. Nothing but chance decided which group each subject landed in, so any systematic difference in the outcome has nowhere else to come from.
When the data was only observed, not assigned β a dashboard, a survey, a historical dataset β a correlation is worth acting on only after checking the alternatives Hill's viewpoints raise: is the association bigger than the noise in the data; does it show up again in a different population or period; did the proposed cause happen before the proposed effect, not after; and is there a specific, nameable third variable that could produce both sides on its own? That check matters more, not less, when the correlation already confirms what you suspected β that pull is confirmation bias, and it is exactly when scrutiny usually drops. When the honest answer to the confounder question is "maybe, nobody has checked," the correlation is a hypothesis worth testing, not a decision worth making.
082 min
Where This Goes Wrong
Three distinct mechanisms produce a correlation with no causal link behind it, and each has its own tell.
A confounding variable is a hidden third cause behind both halves of the correlation. The textbook case: ice cream sales and drowning deaths rise and fall together across the year, but neither causes the other. Ice cream is sold during the hot summer months at a much greater rate than during colder times, and it is during these hot summer months that people are more likely to engage in activities involving water, such as swimming β and swimming, not ice cream, is what creates drowning risk. The tell is a third, unmeasured variable that plausibly moves both halves of the pair.
Reverse causation runs the arrow backward from how it is first read. Windmill rotation and wind speed correlate closely, but wind can be observed in places where there are no windmills or non-rotating windmills, and there are good reasons to believe that wind existed before the invention of windmills. The tell is asking which variable could exist unchanged without the other β that one is usually the cause.
Coincidence needs no mechanism at all. Test enough variable pairs β the earthquake example above came from a project that ran hundreds of millions of comparisons β and some will correlate by chance, at a strength that would look meaningful in any single report. The tell is not in the data; it is asking how many other pairings were tested before this one got reported.
091 min
Correlation and Causation vs. Nearby Concepts
The concept most often collapsed into this one is the confounding variable itself β the specific hidden third cause, rather than the general warning that one might exist. "Correlation is not causation" is the warning that something else might explain an association. "Confounding variable" is the name for that something else, once it has actually been found and named. A confounding variable is simply a hidden third variable that affects both of the variables observed to be correlated β with the distinction that, unlike a merely hidden variable, a confounding variable need not be hidden and may thus be corrected for in an analysis. Once it is named, a researcher can measure it directly and control for it, which turns an unresolved correlation into a causal claim that has actually been tested.
?7 questions
Questions people ask
What is the difference between correlation and causation?
How do you calculate correlation?
What is a good example of correlation without causation?
How do you avoid mistaking correlation for causation?
What is a confounding variable?
Does statistical significance prove causation?
Can an A/B test prove causation where a correlation cannot?
βΆ3 videos
Watch
Understanding Correlation and Causation
Correlation vs Causation Explained: Why Patterns Can Mislead Us
Sprouts
What is Correlation & Causation? Correlation vs Causation
Management Tutorials
Β§7 sources
Sources
Mathematical Contributions to the Theory of Evolution, III: Regression, Heredity, and Panmixia β Karl Pearson, 1896
The Environment and Disease: Association or Causation? β Austin Bradford Hill, 1965
Why do we Sometimes get Nonsense-Correlations between Time-Series? β G. Udny Yule, 1926
Bradford Hill criteria β Wikipedia
Show all 7 sourcesShow fewer sources
Correlation does not imply causation β Wikipedia
The number of movies Nicolas Cage appeared in correlates with Magnitude 7.0-7.9 Earthquakes in the US β Tyler Vigen, Spurious Correlations
Correlation vs. Causation β Jim Frost, Statistics By Jim (secondary; page blocks automated requests)


