Wang et al. 2024/2025's ten-generation iterated-retraining experiment shows political-bias amplification persists independently of model collapse
Wang, Wu, Zhang, Guan, Jain, Lu, Gupta & Koshiyama (Holistic AI, with UCL, Emory, and University of Maryland co-authors) ran a genuine ten-generation iterated-fine-tuning chain on GPT-2, following the design Shumailov et al. established for general model collapse: "G0 then generates a synthetic dataset, D0 ... This dataset D0 is used to fine-tune the Generation 1 (G1) model ... The process continues up to Generation 10 (G10), where each Gi model is fine-tuned on the synthetic data Di−1 produced by model Gi − 1." Unlike the general model-collapse literature, this design tracks a specific bias — political-ideology lean in sentence-continuation on U.S. news text — across that chain, using a validated benchmark "specifically designed to measure political bias amplification in LLMs."
The paper's headline finding: "bias amplification persists independently of model collapse, even when the latter is effectively controlled," with a companion mechanistic result that "largely distinct neuron populations" drive the two effects. This is an empirical, not merely theoretical, demonstration that a specific bias can compound across training generations through a mechanism separable from general distributional degradation.
No citation, reference-selection, or bibliometric variable appears in the design — the bias tracked is political, not citation-popularity — so this does not itself test whether citation-popularity bias compounds across LLM training generations. It establishes instead that the experimental design the question calls for (multi-generation iterated retraining, isolating one bias's trajectory from general model-collapse via a validated benchmark and neuron-level analysis) is buildable and has already produced a clean positive result for a structurally similar bias.
Source
“we perform iterative fine-tuning. First, GPT-2 is fine-tuned on the 1,518 real news articles ... to yield the Generation 0 (G0) model. G0 then generates a synthetic dataset, D0 ... This dataset D0 is used to fine-tune the Generation 1 (G1) model ... The process continues up to Generation 10 (G10), where each Gi model is fine-tuned on the synthetic data Di−1 produced by model Gi − 1.”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-09-14-has-any-study-run-a-genuine-multi-generation.md, 2026-09-14 (headless) · raw markdown