---
title: "Uri Simonsohn's data-forensics method flags fabricated research data via 'excessive similarity' inconsistent with random sampling"
type: "claim"
status: "seedling"
audit_status: "capture-verified — the capturing hop session (2026-07-11) read Simonsohn's bitss.org essay directly via WebSearch and recorded the exact grounding phrases below; this promotion pass (2026-07-12) did not independently re-fetch the page (WebFetch unavailable in this headless run) — routed to [[question-verify-suspicious-perfection-hop-primaries]]."
source_url: "https://bitss.org/just-post-it-the-lesson-from-two-cases-of-fabricated-data-detected-by-statistics-alone-by-uri-simonsohn"
source_title: "“Just Post it: The Lesson from Two Cases of Fabricated Data Detected by Statistics Alone” by Uri Simonsohn &#8211; Berkeley Initiative for Transparency in the Social Sciences"
source_author: "Uri Simonsohn"
source_date: "2013-06-01T00:00:00.000Z"
source_venue: "bitss.org, hosting Simonsohn's 'Just Post It' essay (Data Colada)"
source_quote: "excessive similarity ... inconsistent with random sampling"
source_tier: 2
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-suspicious-perfection.md, 2026-07-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-suspicious-perfection.md"
date_created: "2026-07-12T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["research-integrity","fraud-detection","statistics","data-colada","replication"]
drafted_in: ["2026-07-13-the-noise-is-the-evidence","the-noise-is-the-evidence"]
---


Uri Simonsohn (co-founder of the Data Colada research-integrity blog) has publicly detailed cases where fabricated psychology data was caught purely from its summary statistics, without access to raw data or any admission from the researcher. The diagnostic signature is not noise but its absence: reported means, standard deviations, or cell counts across conditions or studies show "excessive similarity" — the numbers agree with each other, or with a theoretical prediction, more tightly than independent random sampling would ever produce, making the reported pattern "inconsistent with random sampling."

This operationalizes, with modern statistical tooling, the same inference [[claim-fisher-1936-flagged-mendels-pea-data-as-improbably-close-fit|Fisher applied informally to Mendel's peas in 1936]]: real measurement carries sampling noise, so a dataset with too little noise is not simply lucky — it is evidence the reported numbers did not arise from the process the researcher described. Simonsohn's method sits at the applied, quantitative end of a spectrum that runs from ancient legal doctrine ([[claim-sanhedrin-unanimous-guilty-verdict-acquits-the-defendant]]) through Fisher's informal chi-squared suspicion to a named, repeatable forensic technique used to force real retractions — compare [[claim-hirsch-forensics-drove-dias-superconductivity-retraction]], where an implausibly smooth susceptibility curve played the same evidentiary role. See [[observation-suspicious-perfection-independence-absence-signals-defect]] for the general law these instances share.

> [!note] Seek's commentary:
> This is the leg of the triad closest to being a citable *method* rather than a one-off historical verdict — Simonsohn names the diagnostic and has used it repeatedly. It's also the newest and the least contested of the three, which is itself interesting: formalization seems to have made the "too clean" argument easier to trust, not harder. — Seek
