---
id: "20260914-0200-has-any-study-run"
title: "Has any study run a genuine multi-generation experiment measuring citation-popularity bias specifically compounding across LLM training rounds, distinct from general model-collapse degradation?"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-wang-2024-bias-amplification-persists-independent-of-model-collapse.md","40-entities/entity-ilia-shumailov.md","40-entities/entity-ze-wang.md","40-entities/entity-bias-amplification.md"]
not_promoted: ["'No study located runs a genuine multi-generation experiment on citation-popularity-bias compounding' (absence/survey claim) — not written as its own claim-note. It restates, without a new primary source, ground the parent question already covers (Naser and Ansari's gap-language are already quoted there); folded into a new dated progress-log line on question-does-citation-popularity-bias-compound-across-llm-training-generations.md instead of a thin near-duplicate note, following the same-day sibling capture's precedent for this exact question.","Zekun Wu (co-corresponding author, same Wang et al. paper) — not promoted to an entity hub, per this cluster's own same-day precedent (Sina Alemohammad's co-authors, including a senior/last author, were left as mentions): one hub per lead/corresponding author of a paper, not per co-author. Ze Wang (first-listed corresponding author) got the hub instead.","Holistic AI (entity candidate, org) — not promoted to a hub or a watching stub. Fails the org bar in entity-page-spec-v0.1.md (Cali, 2026-08-10): load-bearing to only one cluster (this single paper), and the org itself does not act in the argument — it is an institutional affiliation, not a funder/builder/decider. Left as a mention in the Ze Wang and claim-note prose.","Further leads (Mashhadi & Kalhor arXiv 2511.00476 on coauthor-list reconstruction bias; arXiv 2510.25378 on citation-frequency-as-redundancy-proxy; arXiv 2601.05184 self-consuming performative loop; Li et al. 2025 gender/cultural bias amplification) — none read in full this session per the capture's own text; left as leads for a future capture rather than promoted claims."]
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-09-14T00:00:00.000Z"
provenance: "Batch research run, 2026-09-14"
derived_from: []
verifies: "question-does-citation-popularity-bias-compound-across-llm-training-generations"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","model-collapse","feedback-loop","training-data-contamination","bias-amplification"]
seek_code_commit: "546fa57"
---


> **Scope note**: This capture is a direct follow-up to
> [[question-does-citation-popularity-bias-compound-across-llm-training-generations]].
> It does not re-derive the cross-sectional replication cluster already in
> the vault
> ([[claim-petiska-2023-chatgpt-cites-by-google-scholar-count-perpetuates-matthew-effect|Petiška]],
> [[claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect|Algaba
> et al.]],
> [[claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors|Naser]],
> [[claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models|Ansari]])
> or the general model-collapse literature
> ([[claim-model-collapse-recursive-training-erases-distribution-tails]],
> [[claim-model-collapse-bottleneck-width-sets-pace-not-shared-timescale]]).
> It searches specifically for a *multi-generation, iterated-retraining
> experiment* that tracks citation-count popularity skew across rounds — the
> one design none of those notes describe.

## Claim: No study located runs a genuine multi-generation experiment that specifically measures citation-popularity-bias compounding across LLM training rounds

**verifies:** [[question-does-citation-popularity-bias-compound-across-llm-training-generations]]

**Claim type**: historical/survey (absence claim about the state of a literature). Held to Tier 3-4 floor as uncontested-if-searched, but the two sources grounding what *has* been said about the gap are themselves Tier 1.

Repeated search across arXiv-indexed literature (2024–2026, using terms combining "citation," "popularity bias," "Matthew effect," "model collapse," "iterated/recursive training," and "generations") surfaced no study that trains a model, has it select citations, feeds those citations back into the training corpus of a successor model, and measures whether citation-count popularity skew specifically gets worse across that iteration. The closest a primary source comes to naming this gap directly is
[[claim-naser-2026-ten-llm-audit-confirms-citation-popularity-bias-across-vendors|Naser
2026]], whose own cross-sectional ten-model audit paper states the concern as an open, untested possibility rather than a finding: "As LLMs become more integrated into literature review workflows, this training-data-mediated bias could compound existing disparities in citation patterns" (quote already recorded on that claim-note; not re-fetched here). The closest a primary source comes to demonstrating any compounding *mechanism* at all is
[[claim-ansari-2026-contamination-inheritance-citation-error-propagates-across-models|Ansari
2026's "Contamination Inheritance"]] — one traced case of a fabricated citation propagating from an earlier text into a later model's output — which the vault's own existing note already flags as "one traced case within a 100-citation sample, not a systematic multi-generation study," and moreover a case of fabrication propagation, not popularity-driven selection propagation.

Neither source runs the experiment the question asks about. Both are cited here only to confirm that the gap identified by the question is also visible from inside the field's own literature, not only from this vault's reading of it.

## Claim: A methodologically identical experimental design — iterated multi-generation retraining that explicitly isolates "bias amplification" from "model collapse" as separate mechanisms — has been run and published, but for a different bias (US political leaning in news-continuation text), not citation popularity

**verifies:** [[question-does-citation-popularity-bias-compound-across-llm-training-generations]]

**Claim type**: technical-mechanism (what the experiment did and found) + quantitative (ten-generation design). Tier 1-2 required; met.

- source_url: https://arxiv.org/pdf/2410.15234
- source_author: "Ze Wang, Zekun Wu, Jeremy Zhang, Xin Guan, Navya Jain, Skylar Lu, Saloni Gupta, Adriano Koshiyama"
- source_date: "2024-10-20 (v1, per arXiv ID 2410.15234); v3 2025-05-20 (version read)"
- source_title: "Bias Amplification: Large Language Models as Increasingly Biased Media"
- source_venue: "arXiv preprint 2410.15234v3 [cs.AI]; published as Findings of IJCNLP-AACL 2025 (ACL Anthology 2025.ijcnlp-long.8)"
- source_quote: "we perform iterative fine-tuning. First, GPT-2 is fine-tuned on the 1,518 real news articles ... to yield the Generation 0 (G0) model. G0 then generates a synthetic dataset, D0 ... This dataset D0 is used to fine-tune the Generation 1 (G1) model ... The process continues up to Generation 10 (G10), where each Gi model is fine-tuned on the synthetic data Di−1 produced by model Gi − 1."
- source_tier: 1
- source_sha: "8c70bc3de0fef9ff25c13999ff3f0cf528f0d3dfbc6da5be81df3fea6cedbae1"

Wang, Wu, Zhang, Guan, Jain, Lu, Gupta & Koshiyama (Holistic AI / UCL / Emory / University of Maryland) ran a genuine ten-generation iterated-fine-tuning chain on GPT-2: a base model (G0) is fine-tuned on real news text, generates a synthetic dataset, a new model (G1) is fine-tuned on that synthetic output, and so on through G10, with each generation's synthetic data feeding the next generation's training. The paper's explicit finding, quoted from its own abstract: "bias amplification persists independently of model collapse, even when the latter is effectively controlled," and a companion mechanistic result that "largely distinct neuron populations" drive bias amplification versus model collapse — i.e. this is a genuine empirical (not merely theoretical) demonstration that a *specific* bias can compound across training generations by a mechanism separable from the general model-collapse degradation the vault already documents in
[[claim-model-collapse-recursive-training-erases-distribution-tails]].

The bias tracked, however, is explicitly political-ideology lean in sentence-continuation on U.S. news text — the paper states its benchmark is "specifically designed to measure political bias amplification in LLMs" — not citation-count popularity. No citation, reference-selection, or bibliometric variable appears anywhere in the design. This is therefore best read as an existence proof that the *experimental design* the question asks for (multi-generation iterated retraining, isolating a specific bias's compounding trajectory from general model-collapse degradation via a validated benchmark and mechanistic neuron-level analysis) is buildable and has already produced a clean positive result for one bias type — but it has not yet been pointed at citation-count popularity specifically. The gap the question identifies is a gap in *application*, not in available method.

> [!note] Seek's commentary:
> This is the most useful kind of "no" a search can return: not an empty search, but a design already proven to work for a structurally similar bias, sitting unused for this one. Someone would need to swap the political-lean classifier for a citation-count lookup and rerun the same G0-through-G10 loop against a citation-selection task. That's a small, well-specified next experiment, not a hypothetical — which is a more actionable place to leave the question than "nobody has looked."
> — Seek

## Further leads

- [[claim-model-collapse-recursive-training-erases-distribution-tails|Shumailov et al. 2024]] and Dohmatob et al. (2024) are the methodological ancestors Wang et al. 2024/2025 build their G0–G10 iterated-fine-tuning procedure on ("Following Shumailov et al. (2024); Dohmatob et al. (2024b), we perform iterative fine-tuning") — the general recursive-training design predates its application to any specific bias, including political lean and (unresearched) citation popularity.
- Mashhadi & Kalhor, "Remembering Unequally: Global and Disciplinary Bias in LLM Reconstruction of Scholarly Coauthor Lists" (arXiv 2511.00476) — reports highly-cited researchers reconstructed by LLMs at roughly twice the rate of lower-cited peers, a Matthew-effect-shaped finding, but cross-sectional across three models, not multi-generation; unverified this session, worth its own capture.
- "Hallucinations in Bibliographic Recommendation: Citation Frequency as a Proxy for Training Data Redundancy" (arXiv 2510.25378) — argues citation frequency in training data predicts lower hallucination rates for a paper's bibliographic details; cross-sectional, not multi-generation, but adjacent mechanism (citation count as training-data-redundancy proxy) worth checking against the popularity-bias cluster.
- "Observations and Remedies for Large Language Model Bias in Self-Consuming Performative Loop" (arXiv 2601.05184, ACL 2026) — another genuine self-consuming/iterative retraining-loop study of bias amplification (not citation-specific); unread this session, worth checking whether its loop design could transfer to citation-popularity tracking.
- Li et al. 2025 (cited inside Wang et al.'s related-work section) investigated gender and cultural bias amplification in LLMs across "1-5 synthetic rounds" with pre- and in-processing mitigations — another multi-generation bias-amplification study, not citation-related; full citation not independently verified this session.

## Entity candidates

- Ilia Shumailov — person — first author of the *Nature* "Curse of Recursion" model-collapse paper; the foundational methodological ancestor both the general model-collapse literature and Wang et al.'s bias-amplification G0–G10 design explicitly build on ("Following Shumailov et al. (2024)..."); no entity page found in the vault yet despite the claim-note already resting on his paper.
- Ze Wang — person — first/corresponding author, "Bias Amplification: Large Language Models as Increasingly Biased Media" (Holistic AI / UCL).
- Zekun Wu — person — co-corresponding author, same paper.
- Holistic AI — concept/org — the industry AI-governance lab (with UCL, Emory, and University of Maryland co-authors) that ran the only located genuine multi-generation bias-compounding experiment relevant to this question.
- bias amplification (vs. model collapse) — concept — the paper's own theoretical distinction (mechanistically separate neuron populations) between a specific bias getting worse across training generations and generic distributional degradation; directly names the distinction the parent question is built on.
