The hot-hand "fallacy" was the fallacy: a famous null is a measurement artifact
The famous 1985 'no hot hand' result is an artifact of its own method: GVT's streak estimator reads −7.9pp even on a shooter with no hot hand. The reversal is Miller & Sanjurjo's (2018); we add a runnable null that reproduces it. Measured, with model and falsifier.
The claim. In 1985, Gilovich, Vallone & Tversky concluded that the basketball "hot hand" is a cognitive illusion: conditioning on a streak of made shots does not raise the probability of the next make. It became the textbook example of humans seeing patterns in randomness — cited for decades as settled.
What we measured. We took GVT's exact estimator — P(hit | 3 previous hits) − P(hit | 3 previous misses), computed per player as they did — and ran it on a shooter with no hot hand by construction: independent shots at a fixed 50% rate, ~6,000 records of 100 shots. If the method were unbiased it should return ~0. It returns −7.9 percentage points. (The bias is an exact expected value — it generalizes Miller & Sanjurjo's 5/12 coin example; the t = −27.7 only reflects Monte-Carlo precision across 6,000 records, not a per-player significance test.) It scales with streak length — −3.3pp at streak 2, −17pp at streak 4 — shrinks as records lengthen (~−3pp at 250 shots), and is robust to the base rate (−8.2pp at a 46% shooter).
| streak length | estimator on a TRUE no-hot-hand shooter | (unbiased would be ~0) |
|---|---|---|
| 2 | −3.3pp (t=−17) | biased |
| 3 | −7.9pp (t=−28) | biased |
| 4 | −17.0pp (t=−39) | biased |
Why it happens. This is the Miller–Sanjurjo selection effect (2018): in any finite sequence, the shots that immediately follow a run of hits are, on average, drawn from a slightly hit-depleted remainder of the sequence. So the "after a streak" sample is mechanically biased downward — before a single real player is observed. The estimator measures its own selection bias, not the player. The bias lives in GVT's per-player averaging over finite records: pool every shot from every player together and it disappears — but GVT, like most streak analyses, averaged within each ~100-shot record, which is exactly where Miller & Sanjurjo showed the selection effect bites.
What this means. GVT's headline streak result was not a measurement of players; it is the signature of a biased estimator applied to a random process. A genuinely streaky shooter would have to overcome a built-in headwind of ~8 points (at GVT's 100-shot record length) just to register as 'no effect' — so their canonical evidence against the hot hand is consistent with a real hot hand having been masked. This reversal is Miller & Sanjurjo's published result (Econometrica, 2018), not ours; our only contribution is the runnable null that reproduces their bias from first principles. (GVT's broader claim — that people overestimate streakiness — is a separate question this doesn't settle.)
FalsifierIf the GVT estimator returned ~0 on a constructed independent shooter, the method would be unbiased and this critique would be void. It does not. The correct test of the original question is to apply the bias-corrected estimator to real shooting logs; if that still shows no effect, GVT's conclusion is restored on sound footing. (Bias-corrected re-analyses of GVT's own controlled-shooting data find a substantial hot hand — about +11 to +13 points, Miller & Sanjurjo; in-game effects are smaller and still debated.)
Why we publish it. This is exactly what a replication ledger is for: a number nearly everyone trusts that a clean, transparent null model reproduces from first principles. We post the model so anyone can run it.
FAQ
Is the basketball hot hand really a myth? The famous 1985 'fallacy' rested on a biased estimator: GVT's streak measure reads negative even on a shooter with no hot hand, so it manufactured a null. Bias-correct it (Miller & Sanjurjo 2018) and GVT's own controlled-shooting data show a real, sizable hot hand (~+11–13pp). Their separate point — that fans overestimate streakiness — isn't settled by this.
How big is the bias in the original method? On a truly no-hot-hand shooter (where an unbiased estimator would read ~0), the streak estimator reads −3.3pp at streak length 2 (t = −17), −7.9pp at length 3 (t = −28), and −17.0pp at length 4 (t = −39). The longer the streak conditioned on, the worse the artifact.
Where does the bias come from? From selection: conditioning on a finite streak of makes changes the sampling of the very next shot, pulling the estimate below the true rate. It is the same finite-sample selection effect later identified by Miller & Sanjurjo.
What is the lesson beyond basketball? A famous null can be a property of the estimator, not the world. Before believing “no effect,” test the estimator on a ground-truth case where the answer is known — exactly what we did here.
References
Gilovich, Vallone & Tversky (1985), "The Hot Hand in Basketball", Cognitive Psychology 17:295–314. Miller & Sanjurjo (2018), "Surprised by the Hot Hand Fallacy? A Truth in the Law of Small Numbers", Econometrica 86:2019–2047 (working paper 2015).