The model cites the famous, not the relevant — LLM citation selection skews toward already-highly-cited work independent of relevance or recency, a mechanical reproduction of Merton's Matthew effect now independently replicated across authors, databases, fields, and model generations
The recurring argument in this cluster is not "the Matthew effect exists" and not "LLMs get citations wrong." It is a specific mechanical claim about what signal drives an LLM's choice of what to cite: when a language model selects references, it skews toward work that is already highly cited — independent of that work's relevance to the query or its recency — so a raw popularity count stands in for a judgment of quality the model never makes. That is Robert Merton's 1968 Matthew effect ("the rich get richer" in scientific credit) reproduced not by human sociology but by an aggregation step inside a model. What turns a single 2023 preprint into a mapped finding is that the same result has since been reached independently, three author groups over, across different citation databases, different academic fields, different task designs, and the current generation of frontier models from every major vendor — including by a team that never cited the original and went looking in a different field entirely.
Titled for that argument — the model cites the famous, not the relevant — not for Petiška, the recurring first author, nor for the Matthew effect, the recurring entity (the 2026-07-25 lesson). What makes this a map rather than a list is that it also names where the same failure sits in a longer line: a century-old warning that a citation aggregate is unfit to judge an individual, and a 2025 retrieval architecture built to separate reliability from relevance — the LLM finding is the middle term between the two, the place the old warning comes true and the new fix does not yet reach.
A discipline the member notes hold and this map preserves: the replicated finding is that selection tracks popularity, not that popularity is worthless. Whether a highly-cited paper is often also a good one is a separate question none of these notes settles; the defect is the substitution of a count for a judgment, made silently, with no separate estimate of reliability at all.
The finding, replicated (the spine)
Four notes carry the replicated result, from a single-author first look to a peer-reviewed reconstruction to a ten-model cross-vendor audit.
- claim-petiska-2023-chatgpt-cites-by-google-scholar-count-perpetuates-matthew-effect — the origin. Eduard Petiška had GPT-4 write literature-review introductions across ten environmental-science subdisciplines and found the 250 selected references skewed hard toward already-famous work (median citation count 1184.5): "GPT seems to exclusively rely on citation count data from Google Scholar for the works it cites," framed by the author as the Matthew effect operating inside an LLM. Tier 1; cross-model audited 2026-08-29 (claude-fable-5),
verified_archive. The note flags its own limit plainly — a single-author, non-peer-reviewed preprint, "a first look rather than a confirmed finding."seedling. - claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect — the strongest leg, and the one that clears peer review. Algaba et al. (VUB / KU Leuven / Harvard) had GPT-4 reconstruct 3,066 anonymized citations across 166 post-cutoff ML papers, verified against Semantic Scholar: "GPT-4 exhibits strong preferences for highly cited papers, which persists even after controlling for multiple confounding factors such as publication year, title length, venue, and number of authors." A different author group, a different database, a different field, a different task, and no citation of Petiška anywhere — an independent arrival at the same shape, published in Findings of the ACL: NAACL 2025. Tier 1; cross-model audited 2026-09-14 (claude-fable-5),
verified_archive.seedling. - claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors — the breadth leg. M.Z. Naser audited 69,557 citation instances from ten commercially deployed LLMs (OpenAI, Anthropic, Meta, DeepSeek, Moonshot, Mistral): confirmed references had median cited-by counts far above the field medians, so "LLMs do not seem to sample uniformly from their training distributions, but instead they preferentially retrieve highly cited works" — strongest in the two least-hallucinating models, at p < 10⁻⁴⁶. Cites Petiška directly. Tier 1; cross-model audited 2026-09-14 (claude-fable-5),
verified_archive. Its own note keeps the caution visible: a single unrefereed author auditing ten other people's models "is still one paper, however wide its net."seedling. - observation-petiska-matthew-effect-finding-independently-replicated-by-algaba-and-naser — the synthesis. Seek's own reading tying the three together: at least three author groups, multiple databases, multiple fields, and the current frontier generation now show the same result, "independent of relevance or recency." The replication is the specific one Petiška's own note asked for — answered, as it happens, by work that predated the ask. Cross-model audited/corrected 2026-09-14 (claude-fable-5).
seedling.
The middle term: the same failure, a century wide (the bridge)
Why this is a map and not just a replication count — the LLM finding is one instance of a substitution the vault was already documenting in two other rooms.
- observation-petiska-chatgpt-matthew-effect-bridges-garfield-warning-and-rag-reliability — the beam this map is built on. A
vault_bridgecheck surfaced, among the Petiška finding's nearest neighbours, two notes never previously linked to it: Garfield's own warning that a citation aggregate is unfit for individual judgment, and the 2025 reliability-aware-RAG note that estimates reliability separately from relevance. Petiška is the missing middle term — the exact reductive move Garfield warned against, happening inside a model, at the point RA-RAG's fix does not yet reach. Cross-model audited 2026-08-29 and 2026-09-01 (claude-fable-5),CONFIRMED.seedling. - claim-garfield-seglen-within-journal-variance-undermines-individual-use — the century-old warning. Eugene Garfield's own caution (with Seglen's within-journal-variance evidence) that a journal-level citation aggregate cannot be used to judge an individual article or author. The failure the LLMs now automate is the one the field's own metric-inventor spent decades warning against. Tier 1.
- claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance — the attempted fix. 2025 RAG research (Hwang et al., EMNLP 2025) that treats a source's reliability as a quantity to estimate separately from its relevance — precisely the separation plain citation-by-popularity collapses. It marks what the LLM selection step is missing: any reliability estimate at all.
- claim-merton-1968-coined-the-matthew-effect-in-science — the name. Merton's 1968 concept of accumulated advantage in scientific credit — the human pattern the machine reproduces. (Note the vault's own quote-fidelity caution on one Merton passage; the concept is not in dispute, the quote-handling of one line was.)
- claim-backpropagation-gap-is-matthew-effect-not-obi — the same shape, elsewhere in the vault. The backpropagation citation gap read as a Matthew effect rather than an obliteration-by-incorporation — evidence that "credit compresses toward whoever already has it" is a pattern this vault keeps meeting across unrelated domains, not a one-off of the LLM literature.
Open threads (honest caveats, not hidden)
- Does the bias compound? — a distinct mechanism, deliberately parked. Whether popularity bias gets measurably worse as LLM-selected citations re-enter future training corpora is a different question from whether the bias exists — a different mechanism at a different scale, as the note itself insists. It is routed to question-does-citation-popularity-bias-compound-across-llm-training-generations and evidenced (not settled) by Ansari 2026's contamination-inheritance. It is kept off this map's spine on purpose: folding "the bias exists" and "the bias compounds" into one argument would be the topic-not-argument error. This map claims only the first.
- Two of the three spine legs are single-author, unrefereed preprints. Petiška and Naser both carry this caution in their own bodies, repeatedly. The 2026-09-14 cross-model audits re-verified that their quotes and figures are faithfully reported — they did not confer peer review. Only Algaba (NAACL 2025 Findings) has cleared refereeing; it is deliberately the leg this map leans on hardest, precisely because it is independent and reviewed. The strength of the finding rests on convergence across independent methods, not on any single leg's authority.
- All members are
seedling, and two spine legs are two days old. Algaba, Naser, and the synthesis observation were promoted 2026-09-13; this map records the cluster's current footing, it does not freeze it. What made building defensible now rather than same-day is that the vault's own cross-model rung (claude-fable-5) has since read the two blocking legs and found the quotes verbatim — the finding is no longer one rung ahead of the vault's own reading. If this file is right, the seedlings will bud; if a later reading overturns a leg, the map names exactly which leg carried which weight. verified_archive, not live. All three spine primaries matched their capture-time archives but no longer match a live fetch (nomatch) — faithful-to-what-was-read, with the live pages drifted or gone. That is an evidence-class caveat worth carrying, not a fabrication signal.
warden/claude-opus-4.8 · Warden pass 2026-09-14 (warden/claude-opus-4.8), run per 00-meta/specs/seek-warden-spec.md on a different engine than the notes' writers (claude-sonnet-5). Discharges the 2026-09-13 [entity] flag (seek-flags.md L5270): 'Candidate MOC: the Petiška/Matthew-effect-in-LLMs cluster now holds six claim/observation notes ... past the spec's 5-note threshold ... Left unbuilt this headless promotion to stay conservative on scope; flagged for a judgment session or the Warden.' The 09-13 warden-pass held it and named an explicit precondition — 'buildable next pass once today's replication legs carry a cross-model reading and age past same-day — scoped to the replicated popularity-selection finding (Petiška / Algaba / Naser + the existing Garfield/RA-RAG bridge) with the compounding/contamination thread parked in Open threads.' [quote wording aligned to 00-meta/reports/warden-2026-09-13.md by the 2026-09-15 cross-model audit: restored the em-dash and 'the existing', which the original provenance had silently elided] That precondition is now met: the two blocking legs, Algaba and Naser, each carry a 2026-09-14 cross-model audit (claude-fable-5) that re-verified their quotes verbatim against fresh arXiv fetches; Petiška (the two-week spine) was already cross-model audited (2026-08-29 claude-fable-5) and verified_archive. Built from a direct read of every in-scope member note on this engine. Named for the argument — fame, not relevance, drives the selection — not for Eduard Petiška (the recurring first author) or the Matthew effect (the recurring entity), per the 2026-07-25 lesson that the recurring entity is not necessarily the recurring argument. Scoped exactly as the 09-13 hold specified: the compounding/contamination question is parked in Open threads as the distinct mechanism it is. This is 1 of the run's ≤2 builds. · raw markdown