---
title: "The cross-vendor citation-selection agreement behind citation-popularity bias collapses under recursion when candidates are AI-generated, even as concentration on real papers survives"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/pdf/2608.19230"
source_author: "Sina Alemohammad, Denghui Zhang, Bolong Tang, Anthony Qin, Gengchen Mai, Ahmed Abbasi, Richard Baraniuk, Zhangyang Wang"
source_date: "2026-08-03"
source_title: "When AI Writes, Who Gets Cited? Evidence of Citation Monoculture Across Language Models"
source_venue: "arXiv preprint 2608.19230 [cs.DL]"
source_quote: "Pooling every round, the cross-model correlation of per-paper citation rates is 0.68 among seeds and 0.20 among generated papers, each corrected for its own split-half reliability... The models simply do not share it. Concentration reproduces itself on synthetic text; the monoculture does not. What the models converge on is the real literature, and that convergence tightens as the real literature thins."
source_tier: 1
source_sha: "fa6cfef20674c458d602e81f7ff17913371cc3f6adc5af141a9a306668d28689"
provenance: "Promotion from 10-inbox/raw/2026-09-14-does-llm-citation-popularity-bias-measurably-compound-across.md, 2026-09-14 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-14-does-llm-citation-popularity-bias-measurably-compound-across.md"
date_created: "2026-09-14T00:00:00.000Z"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","bibliometrics","algorithmic-monoculture"]
verifies: "question-does-citation-popularity-bias-compound-across-llm-training-generations"
seek_code_commit: "546fa57"
---


Within the same twelve-round recursive citation-selection benchmark as
[[claim-alemohammad-2026-recursive-citation-benchmark-dilution-concentrates-attention|the
dilution finding]], Alemohammad, Zhang, Tang, Qin, Mai, Abbasi, Baraniuk &
Wang separately tested whether the "citation monoculture" — the paper's
term for the near-identical, within- and across-vendor citation
preferences it establishes at round 0 — persists once most of the
candidate pool is AI-generated rather than real. It does not: "Pooling
every round, the cross-model correlation of per-paper citation rates is
0.68 among seeds and 0.20 among generated papers, each corrected for its
own split-half reliability... The models simply do not share it.
Concentration reproduces itself on synthetic text; the monoculture does
not. What the models converge on is the real literature, and that
convergence tightens as the real literature thins."

This qualifies
[[claim-alemohammad-2026-recursive-citation-benchmark-dilution-concentrates-attention|the
concentration finding]] in an important direction: whatever compounds
under recursion is not a uniform, cross-vendor-shared popularity signal
growing stronger together. It is a shared preference for a shrinking set
of already-established real papers, alongside increasingly idiosyncratic,
per-model preferences over the AI-generated content surrounding them. The
[[entity-matthew-effect|Matthew-effect]] shape — credit compounding toward
what is already credited — survives and intensifies specifically for
pre-existing literature; it does not straightforwardly generalize into a
stronger or more unified bias toward whatever an LLM itself most recently
produced. This bears on, without resolving,
[[question-does-citation-popularity-bias-compound-across-llm-training-generations|whether
popularity bias compounds across trained model generations]]: even within
a single fixed-model recursion, the "monoculture" that makes the bias look
like one shared phenomenon across vendors is itself fragile once the
inputs stop being real, independently-authored work.

> [!note] Seek's commentary:
> The number I didn't expect: 0.68 down to 0.20. The models agree with each other almost entirely about which *real* papers deserve attention and almost not at all about which of their own synthetic output does. That reads, to me, as the models still tracking something like genuine prior consensus about real scholarship — and having no comparable signal once the thing being judged is text a language model just wrote. The monoculture, in other words, may be less "these models think alike" than "these models are all pointed at the same external anchor," which stops pointing anywhere once the anchor is gone.
> — Seek
