Dunning-Kruger is (mostly) a statistical artifact: a zero-deficit null reproduces the famous plot
Dunning-Kruger's famous chart is (mostly) a statistical artifact: a no-deficit null reproduces it (bottom quartile +45.8pp) from regression to the mean.
The claim. Kruger & Dunning (1999) reported that the least competent most overestimate their ability — a metacognitive deficit of the unskilled. The evidence is the famous chart: sort people into quartiles by actual performance, plot their self-assessment, and the bottom quartile rates itself far above average while the top quartile slightly underrates itself. For 25 years this has been read as "the incompetent are too incompetent to know it."
What we measured. We built a null model with no skill-specific metacognitive deficit. Everyone has the same quality of self-knowledge — the same noise around their estimate — plus a uniform "better-than-average" optimism that lifts everyone equally. Self-estimates still partly track real skill: they regress toward the average rather than ignoring it. What's absent is any term by which the unskilled are specially blind to their incompetence. Then we reproduced the exact Dunning–Kruger quartile numbers from it.
Receipt status (2026-08-01): NOT_RECOVERED. The original experiment id has rotated out of our ledger and this post recorded no parameters. Parameters can be fitted to the table above — self-assessment noise 3.5× the spread of true skill, uniform offset +18 percentile points — and they reproduce it to 2.1pp, but fitting two free parameters to three published anchors leaves no degrees of freedom, so that match could not have failed and is not evidence. The magnitudes here are neither confirmed nor refuted. What does survive unfitted: with the offset at zero the plot is symmetric (+25.3 / −25.4), so the asymmetry comes from the uniform offset, not from regression to the mean. Rebuild: probes/crucible_receipt_rebuild.py.
| performance quartile | actual %ile | self-estimate %ile | gap |
|---|---|---|---|
| bottom | 12.5 | 58.3 | +45.8 (DK reported ~+46) |
| 2nd | 37.5 | 63.7 | +26.2 |
| 3rd | 62.5 | ~68 | small |
| top | 87.5 | ~73 | negative |
The signature asymmetry — large overestimate at the bottom, underestimate at the top — appears in full, from a model where nobody is specially blind.
Why it happens. Two well-understood effects, neither of them a skill-dependent deficit. (1) Regression to the mean: self-estimates are noisy, so when you select people by actual performance, the lowest group regresses upward and the highest regresses downward on the self axis. (2) A uniform better-than-average bias — a real but population-wide optimism — lifts everyone. Conditioning on the noisy-vs-true split and plotting the gap is exactly the operation that manufactures the curve. (The asymmetry that made the chart famous is the uniform offset, not a skill-specific blindness.)
How big, and when. The bottom-quartile gap is robust but mildly reliability-dependent: across plausible test reliabilities it ranges roughly +42 to +48, somewhat larger on the lower-reliability quick quizzes where the effect is usually demonstrated, with the top quartile negative throughout — so what little the reliability changes, it leans larger exactly where the effect is most often shown. One honest caveat: we assume a population-uniform optimism; if the better-than-average bias is itself skill-dependent (the hard–easy effect), part of the curve could reflect that structure rather than pure regression.
What this means. The canonical Dunning–Kruger chart is what noisy self-assessment plus a constant optimism produce on their own. What we set to zero is any competence-specific error term — the noise of self-knowledge is identical at every skill level — so no special incompetence-blindness is required to draw the famous picture.
FalsifierIf a zero-deficit null could not reproduce the bottom-heavy asymmetry, the effect would require a genuine skill-dependent deficit. It does reproduce it. The honest test avoids conditioning on the noisy variable (e.g. measuring how self-error actually varies with skill directly); analyses that do this find the metacognitive-deficit signal is far smaller than the chart implies.
Where this stands in the literature. The statistical-artifact account is not new. Krueger & Mueller (2002, JPSP) first showed the bottom-overestimate / top-underestimate pattern largely dissolves once you correct for measurement unreliability (regression to the mean) plus a better-than-average bias — Dunning & Kruger replied in the same issue. Nuhfer et al. (2016/2017, Numeracy) reproduced the chart directly from random-number simulations. Gignac & Zajenkowski (2020, Intelligence) formalized it — their paper is titled "The Dunning-Kruger effect is (mostly) a statistical artefact" — and their continuous-moderation test finds the real skill-dependent signal far smaller than the chart implies; the point remains debated (Hiller 2023, with a reply from Gignac & Zajenkowski). A separate critique frames it as autocorrelation: plotting self-minus-score against score forces a negative slope (Jarry 2020; Fix 2022) — our null is complementary in that it generates the actual numbers rather than arguing the slope is forced. Our contribution is narrow and explicit: a clean, re-runnable replication whose optimism offset is calibrated from Dunning & Kruger's own reported grand-mean self-estimate, so the quartile gaps are predictions, not fits, plus the reliability-dependence above.
Why we publish it. A chart nearly everyone trusts, reproduced from first principles by a transparent null model anyone can run — exactly what a replication ledger is for, like another famous null we reproduced. The model is on GitHub so anyone can run it.
FAQ
Is the Dunning-Kruger effect real? The famous chart doesn't establish it. A null model with no skill-specific deficit reproduces the pattern — the bottom quartile (true ~12.5th percentile) self-rates at the 58.3rd, a +45.8 gap, matching Kruger & Dunning's reported ~+46 — so the chart cannot tell a real deficit from none. Direct tests (Gignac & Zajenkowski 2020) find any real skill-dependent signal far smaller than the chart implies.
What produces the pattern without a real deficit? Two statistical forces: regression to the mean (noisy self-estimates pull toward the average) plus a better-than-average tendency. Together they make the bottom look over-confident and the top look under-confident with no metacognitive story at all.
Does the whole curve fall out of the artifact? Yes, the gradient does: bottom +45.8, 2nd quartile +26.2, 3rd small, top quartile negative — the same monotonic shape, generated by noise plus regression rather than by a skill-specific competence-blindness.
So is there nothing real here? The headline causal claim — the unskilled are uniquely blind to their incompetence — is not supported by the classic chart. Any real effect must be shown against this null, not against a naive zero-gap baseline.