The Dunning-Kruger effect is autocorrelation

Critics of the famous Dunning–Kruger effect argue that much of its signature pattern—unskilled people overestimating their ability and experts underestimating theirs—can arise purely from statistical artifacts such as regression to the mean, bounded test scores, and how the data are plotted. Others counter that even after accounting for these effects, studies still show a real (though smaller) psychological bias: people tend to rate themselves above average, and self-assessment accuracy improves only weakly with actual skill. The exchange highlights broader concerns about shaky methods in social science, the gap between the original Dunning–Kruger findings and their pop-culture version, and how eager many are to use the effect to explain perceived incompetence in others.

Scope of the Dunning–Kruger (DK) Effect

  • Several commenters note a gap between the original DK finding and the pop-sci meme.
  • Original results:
    • Self-assessed ability correlates positively (not negatively) with actual ability, but weakly.
    • Everyone’s estimates are biased upward (around the ~65th percentile).
    • Lower performers overestimate more; higher performers underestimate somewhat; bias regresses toward a point slightly above average.
  • Pop version (“idiots think they’re geniuses, experts think they’re idiots”) is widely seen as a distortion.

Autocorrelation / Regression-to-the-Mean Critique

  • The linked article simulates random, independent performance and self-assessment.
    • When plotted as “actual vs (self-estimate – actual)”, this automatically yields a DK-like pattern: low scorers overestimate, high scorers underestimate.
  • Many argue this is just regression to the mean / bounded scores, not “autocorrelation” in the standard time-series sense.
  • Point: if random data produce the same shape, then that specific plot is a bad test for a psychological effect.

Rebuttals to the Simulation Argument

  • Others counter that in real life performance and self-assessment are not expected to be independent; a null model of pure randomness is unrealistic.
  • They note differences between DK data and the random simulation:
    • Real self-estimates trend upward with ability (non-flat line).
    • Real estimates cluster above the 50th percentile, unlike uniform random guesses.
  • Therefore, they see evidence for:
    • Systematic overconfidence, and
    • Improved self-assessment with higher skill, though the effect size is small.

Methodological and Statistical Concerns About DK

  • Critiques raised in the thread:
    • Use of bounded percentiles (ceiling/floor effects) and “double dipping” (using test score both for grouping and error) can manufacture the DK curve.
    • Lack of variance/error bars in original plots obscures whether groups truly differ.
    • Tasks include subjective measures (e.g., humor); sample is small, homogeneous (Cornell undergrads seeking extra credit).
    • Replication attempts and later modeling papers suggest much of the canonical DK graph can arise from noise plus bounds, with only a minimal residual effect.

Status, Replication, and Perception

  • Some commenters claim DK has been “debunked” (or heavily qualified) in multiple later papers; others say rebuttals themselves misunderstand the statistics.
  • One recent meta-analysis (cited in the thread) reportedly finds a real but very small DK effect, raising doubts about its practical significance.
  • Several note the social/psychological appeal of DK as a way to label others as “confident idiots,” which may explain its persistence despite methodological controversy.