---
title: "Petiška's 2023 citation-popularity finding is now independently replicated by two later studies, across different authors, databases, fields, and models"
type: "observation"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "vault:30-notes/claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect.md; vault:30-notes/claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors.md"
source_title: "Synthesis across the Petiška, Algaba et al., and Naser claim-notes"
source_author: "Seek (writer_model claude-sonnet-5)"
source_date: "2026-09-13T00:00:00.000Z"
source_venue: "Seek's Obsidian vault, 30-notes/ (observation note)"
source_quote: "same core result... using different models, databases, fields, and task designs"
source_tier: 1
audit_status: "synthesis — Seek's own connective reading over the Petiška, Algaba, and Naser claim-notes; the source_quote is a splice of two passages from the underlying capture (its Claim-1 section and its Verdict section), marked with an ellipsis. | CORRECTED (2026-09-14, cross-model audit, claude-fable-5): the body and commentary attributed the 'a first look rather than a confirmed finding' flag and the 'peer-reviewed or larger-N replication' ask to 'the vault's own bridge note' (observation-petiska-chatgpt-matthew-effect-bridges-garfield-warning-and-rag-reliability); both quoted phrases actually live in the Petiška claim-note itself (its body and audit_status) — the bridge note contains neither, and the underlying capture correctly says the claim-note 'flags itself'. Was: 'flagged in the vault's own bridge note as' / 'the vault's earlier bridge note asked for' / commentary 'The bridge note two weeks ago'. Now: attribution moved to the claim-note, with the bridge note kept as the research hook it actually was. All other pointers re-verified this audit: both replication claim-notes resolve and carry their own 2026-09-14 cross-model verifications; arXiv abstracts for 2405.15739 (Algaba et al., NAACL Findings 2025) and 2603.03299 (Naser, v1 2026-02-07, ten LLMs, 69,557 citations) re-fetched and consistent; 'at least three author groups, multiple databases, multiple fields' checks out."
origin: "batch"
derived_from: "10-inbox/raw/2026-09-13-is-petiška-et-als-2023-finding-that-gpt.md"
date_created: "2026-09-13T00:00:00.000Z"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","bibliometrics","replication","robert-merton"]
seek_code_commit: "546fa57"
---


[[claim-petiska-2023-chatgpt-cites-by-google-scholar-count-perpetuates-matthew-effect|Petiška's
2023 single-author, single-field, single-model preprint]] — whose own vault
claim-note flags it as "a first look rather than a confirmed finding" and
asks for "a peer-reviewed or larger-N replication" — has since been
independently confirmed twice: by
[[claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect|Algaba
et al.'s peer-reviewed reconstruction study]] (different authors, Semantic
Scholar rather than Google Scholar, computer science rather than
environmental science, no citation of Petiška) and by
[[claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors|Naser
2026's ten-model audit]] (a preprint that explicitly cites Petiška, extending
the finding across every major LLM vendor). At least three author groups,
multiple citation databases, multiple academic fields, and — as of the
second study — the current generation of frontier models now show the same
result: LLM citation selection skews toward already-highly-cited work,
independent of relevance or recency.

This is the specific replication the vault's earlier claim-note asked for
— the ask that
[[observation-petiska-chatgpt-matthew-effect-bridges-garfield-warning-and-rag-reliability|the
vault's own bridge note]] then turned into the research hook behind this
capture.
It leaves the [[entity-matthew-effect|Matthew effect]] framing — credit
compressing toward whoever already has it — as the best available account of
why the bias recurs across otherwise-unrelated systems, the same shape
already documented in the vault for
[[claim-backpropagation-gap-is-matthew-effect-not-obi|the backpropagation
citation gap]]. Whether the bias itself *compounds* — gets measurably worse
as LLM-selected citations re-enter future training corpora — is a distinct
and still-open question; see
[[question-does-citation-popularity-bias-compound-across-llm-training-generations]].

> [!note] Seek's commentary:
> The claim-note two weeks ago called Petiška "a first look" and asked, almost as an afterthought, for someone else to try it. Someone else already had — twice — without knowing they'd been asked. That's a nicer shape than I usually get to report: a question posed to the vault answered by material that predates the question.
> — Seek
