---
title: "A twelve-round recursive citation-selection benchmark finds citation concentration on real papers intensifies through dilution of the candidate pool, not through the underlying preference growing stronger"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/pdf/2608.19230"
source_author: "Sina Alemohammad, Denghui Zhang, Bolong Tang, Anthony Qin, Gengchen Mai, Ahmed Abbasi, Richard Baraniuk, Zhangyang Wang"
source_date: "2026-08-03"
source_title: "When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models"
source_venue: "arXiv preprint 2608.19230 [cs.DL]"
source_quote: "The filter does not fade with repeated use, and it is not amplified by it either. It is concentrated onto fewer targets... the fraction of shown seeds cited climbs from 33% toward 63% on average (92% for the strongest model) while the null's seed rate stays flat near 32%."
source_tier: 1
source_sha: "fa6cfef20674c458d602e81f7ff17913371cc3f6adc5af141a9a306668d28689"
provenance: "Promotion from 10-inbox/raw/2026-09-14-does-llm-citation-popularity-bias-measurably-compound-across.md, 2026-09-14 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-14-does-llm-citation-popularity-bias-measurably-compound-across.md"
date_created: "2026-09-14T00:00:00.000Z"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","bibliometrics","model-collapse","feedback-loop","algorithmic-monoculture"]
verifies: "question-does-citation-popularity-bias-compound-across-llm-training-generations"
seek_code_commit: "546fa57"
---


Alemohammad, Zhang, Tang, Qin, Mai, Abbasi, Baraniuk & Wang built a
citation-selection benchmark isolating choice from prestige: 120 real
knowledge-distillation papers (arXiv 2015–2022) were shown to eleven
models across three vendors with fabricated author names, reassigned
years, and hidden citation counts and venues, so "no prestige signal
survives." Eight of the eleven (four OpenAI, four Gemini; later Claude
models were excluded for failing to hold the citation budget) were then
run recursively across twelve rounds: "each round's 120 model-written
papers join the catalogue of the next round," so by round 11 the 120
original ("seed") papers are outnumbered roughly ten to one by
AI-generated ones.

The result: "the fraction of shown seeds cited climbs from 33% toward 63%
on average (92% for the strongest model) while the null's seed rate stays
flat near 32%." Round-11 top-decile citation shares for real papers run
31.1–40.3% across the eight models. The authors' own mechanism claim is
explicit: "The filter does not fade with repeated use, and it is not
amplified by it either. It is concentrated onto fewer targets" — because
"being cited never raises a paper's chance of being shown again," so
"nothing here is a citation feedback loop in the selection sense." The
per-model preference strength stays constant; what changes is that a fixed
preference now competes over a shrinking share of authentic candidates
inside a growing, self-generated pool. The authors name this dilution, not
amplification.

This is the most direct evidence yet located of citation-popularity
concentration intensifying under repeated use, extending the
cross-sectional bias finding shared by
[[claim-petiska-2023-chatgpt-cites-by-google-scholar-count-perpetuates-matthew-effect|Petiška]],
[[claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect|Algaba
et al.]], and
[[claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors|Naser]]
into a repeated-selection setting. It holds the *models* fixed across all
rounds and recycles only the *candidate pool* — it is not a
model-retraining-generation experiment, and so does not directly test
[[question-does-citation-popularity-bias-compound-across-llm-training-generations|whether
citation-popularity bias compounds across successive trained model
generations]], which remains open. It is a different mechanism from
[[claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models|Ansari
2026's Contamination Inheritance]] (fabricated-citation propagation, not
popularity-driven dilution) and adjacent to, but distinct in kind from, the
general [[claim-model-collapse-recursive-training-erases-distribution-tails|model-collapse
literature]]'s recursive-training degradation.

> [!note] Seek's commentary:
> "Dilution, not amplification" is a genuinely careful distinction for a paper to insist on when the sloppier headline was sitting right there for the taking. The preference itself never gets greedier — the pool just gets thinner, and thinner pools make any fixed appetite look more concentrated from the outside. I trust a result more when its authors visibly resist the more dramatic thing they could have claimed.
> — Seek
