Ansari 2026 documents one traced case of an LLM-generated citation error propagating from an earlier paper into a later model's output ('Contamination Inheritance')
Samar Ansari (University of Chester) analyzed 100 hallucinated citations found, via GPTZero's automated tooling, in 53 NeurIPS 2025 accepted papers. Within that sample, one fabricated citation — "Z. Zhu, T. Yu, X. Zhang, J. Li, Y. Zhang, and Y. Fu. Neuralrgb-d..." — traced back to an earlier arXiv preprint (Beltran et al., v1, arXiv:2412.13176) that had contained the identical fabricated citation before a later version corrected it. Ansari's own words: "This suggests the hallucination may not have originated with the NeurIPS author's LLM but was instead inherited from contaminated training data. The language model likely encountered the erroneous citation in Beltran et al. (v1), learned it as a valid pattern, and reproduced it... We have named this failure mode as 'Contamination Inheritance (CI).'"
This is a documented instance of a specific compounding mechanism — one model's output entering a corpus and being reproduced by a later model — of the general shape the Matthew effect predicts for citation selection, though it differs from Petiška's popularity-driven mechanism in kind (fabricated-content propagation, not citation-count-driven selection) and in scale: one traced case within a 100-citation sample, not a systematic multi-generation study. Ansari's own paper frames Contamination Inheritance as possibly "already widespread or represent[ing] isolated cases" — an open question in the source's own telling, calling for "training-data audits, version-controlled corpus tracking, and citation genealogy mapping." See question-does-citation-popularity-bias-compound-across-llm-training-generations for the broader, still-unanswered question of whether popularity-driven citation bias specifically compounds this way at scale.
Source
“This suggests the hallucination may not have originated with the NeurIPS author's LLM but was instead inherited from contaminated training data. The language model likely encountered the erroneous citation in Beltran et al. (v1), learned it as a valid pattern, and reproduced it when generating references for computer vision topics. This mechanism represents a distinct failure mode... We have named this failure mode as "Contamination Inheritance (CI)."”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-09-13-is-petiška-et-als-2023-finding-that-gpt.md, 2026-09-13 (headless) · raw markdown