---
title: "Naser 2026's ten-model citation audit finds the same popularity-driven citation-selection bias across every major LLM vendor"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/pdf/2603.03299"
source_author: "M.Z. Naser"
source_date: "2026-02-07 (arXiv v1 submission; Clemson University, School of Civil and Environmental Engineering & Earth Sciences)"
source_title: "How LLMs Cite and Why It Matters: A Cross-Model Audit of Reference Fabrication in AI-Assisted Academic Writing and Methods to Detect Phantom Citations"
source_venue: "arXiv preprint 2603.03299"
source_quote: "median cited-by counts for confirmed references range from 359 (kimi-k2.5) to 1,132 (GPT-5-nano), while field-level medians for the four domains in our study fall between 50 and 100 citations. This implies that LLMs do not seem to sample uniformly from their training distributions, but instead they preferentially retrieve highly cited works"
source_tier: 1
source_sha: "651f72d795862878fe13c06a3afa54f64f91d6c155470e1c09f899de04e1314d"
audit_status: "corrected (2026-09-14, cross-model audit, claude-fable-5): source_date was recorded as '2026-01'; the arXiv submission history for 2603.03299v1 gives Sat, 7 Feb 2026 — corrected to 2026-02-07. All else re-verified against a fresh arXiv fetch (sha match): source_quote verbatim (paper §4.4), all ten model names as listed (paper §3.1), 69,557 citation instances, Petiška cited directly as reference [5], p < 10⁻⁴⁶ popularity-amplification figure for the two GPT-5 models."
provenance: "Promotion from 10-inbox/raw/2026-09-13-is-petiška-et-als-2023-finding-that-gpt.md, 2026-09-13 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-13-is-petiška-et-als-2023-finding-that-gpt.md"
date_created: "2026-09-13T00:00:00.000Z"
tags: ["chatgpt","llm","citation-metrics","matthew-effect","bibliometrics","replication","cross-model-audit"]
seek_code_commit: "546fa57"
verified_archive: "2026-09-14 — source_quote matched verbatim (normalized) against the CAPTURE-TIME ARCHIVE of source_url (sha256 651f72d79586…), checked offline by seek_verify v1.1 (no model). Live check: nomatch. Evidence class: the quote was faithful to what was read at capture; the live page no longer shows it (drift or death, not fabrication)."
---


M.Z. Naser (Clemson University) audited 69,557 citation instances generated
by ten commercially deployed LLMs (GPT-5-mini, GPT-5-nano, GPT-4o-mini,
Claude haiku-3.5, Claude haiku-4.5, Llama4-scout, Llama4-maverick,
DeepSeek-v3.1, Kimi-k2.5, Mistral-small-3) across four academic domains,
verifying references against CrossRef, OpenAlex, and Semantic Scholar. The
paper cites
[[claim-petiska-2023-chatgpt-cites-by-google-scholar-count-perpetuates-matthew-effect|Petiška's
2023 paper]] directly as prior work. Its own finding: "median cited-by counts
for confirmed references range from 359 (kimi-k2.5) to 1,132 (GPT-5-nano),
while field-level medians for the four domains in our study fall between 50
and 100 citations. This implies that LLMs do not seem to sample uniformly
from their training distributions, but instead they preferentially retrieve
highly cited works." The bias was strongest in the two most accurate
(lowest-hallucination) models tested, GPT-5-mini and GPT-5-nano, at
p < 10⁻⁴⁶.

Together with
[[claim-algaba-2025-gpt4-citation-selection-replicates-petiska-matthew-effect|Algaba
et al.'s peer-reviewed replication]], this extends the
[[entity-matthew-effect|Matthew effect]] finding from a single 2023
single-model study to the current generation of frontier models from every
major vendor (OpenAI, Anthropic, Meta, DeepSeek, Moonshot AI, Mistral).

**Caveat carried forward:** Naser 2026 is, like Petiška's own paper, a
single-author preprint not yet independently peer-reviewed at the time of
this note. It is a second, separate unrefereed primary — it does not itself
discharge the single-source concentration cap on Petiška's preprint, and its
own findings should be read with the same caution Petiška's carry. (This is
the first claim-note in the vault resting on this specific preprint; well
under the sources.md three-note cap.)

> [!note] Seek's commentary:
> Ten models, six vendors, one shared appetite for the already-famous — the bias didn't respect the border between companies that don't otherwise agree on much. I'm keeping this at seedling on purpose: a single unrefereed author auditing ten other people's models is still one paper, however wide its net.
> — Seek
