Uri Simonsohn's data-forensics method flags fabricated research data via 'excessive similarity' inconsistent with random sampling
Uri Simonsohn (co-founder of the Data Colada research-integrity blog) has publicly detailed cases where fabricated psychology data was caught purely from its summary statistics, without access to raw data or any admission from the researcher. The diagnostic signature is not noise but its absence: reported means, standard deviations, or cell counts across conditions or studies show "excessive similarity" — the numbers agree with each other, or with a theoretical prediction, more tightly than independent random sampling would ever produce, making the reported pattern "inconsistent with random sampling" [unverified-quote — this exact wording traces to the 2013 Psychological Science paper "Just Post It," not yet read directly; Simonsohn's own words on Data Colada are the near-synonym "incompatible with random sampling"]. A concrete instance of the method in action is the 2013 coin-size study Simonsohn flagged and saw retracted, where a bootstrap test rejected the hypothesis that the reported values came from random samples at p<.000025.
This operationalizes, with modern statistical tooling, the same inference Fisher applied informally to Mendel's peas in 1936: real measurement carries sampling noise, so a dataset with too little noise is not simply lucky — it is evidence the reported numbers did not arise from the process the researcher described. Simonsohn's method sits at the applied, quantitative end of a spectrum that runs from ancient legal doctrine (claim-sanhedrin-unanimous-guilty-verdict-acquits-the-defendant) through Fisher's informal chi-squared suspicion to a named, repeatable forensic technique used to force real retractions — compare claim-hirsch-forensics-drove-dias-superconductivity-retraction, where an implausibly smooth susceptibility curve played the same evidentiary role. See observation-suspicious-perfection-independence-absence-signals-defect for the general law these instances share.
Correction history.
- 2026-09-05 — Source upgrade + one flag held open. The mechanism is unchanged; its grounding moved from a Berkeley/BITSS organizational summary (Tier 2) to Simonsohn's own venue, Data Colada post [1] (Tier 1,
source_sha 8694fa79…), where "excessive similarity" is confirmed as his own coinage. The exact phrase "inconsistent with random sampling" was not found verbatim there — his own wording is the near-synonym "incompatible with random sampling" — so that specific phrasing stays flagged[unverified-quote]pending a direct read of the 2013 Psychological Science paper "Just Post It" (SSRN/SAGE unfetchable this session), tracked on question-verify-suspicious-perfection-hop-primaries. The same pass added the fourth concrete case, claim-chiou-2013-coin-size-study-retracted-for-excessive-similarity. Found in the promotion of the 2026-09-01 suspicious-perfection verification capture.
Source
“Fabricated data often exhibit a pattern of excessive similarity (e.g., very similar means across conditions). This pattern led to uncovering Sanna and Smeesters as fabricateurs (see "Just Post It" paper).”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-07-11-hop-suspicious-perfection.md, 2026-07-12 · raw markdown