What is AuthorityBench's actual identity — arXiv id, venue, authors — and does the Qwen3-32B ρ 75.28% DomainAuth figure check out?
claim-document-text-degrades-llm-source-authority-judgment rests on a source whose provenance does not internally cohere, and the note carries an [unverified-quant] flag until this resolves.
The inconsistency. The promoting capture (10-inbox/raw/2026-07-11-hop-admiralty-code-to-rag.md) calls this "2025 RAG research" but cites a 2026 arXiv HTML id, https://arxiv.org/html/2603.25092 (arXiv 26MM. = 2026). It separately lists an ACL-Anthology URL, https://aclanthology.org/2025.emnlp-main.1738/ (EMNLP 2025), for the companion RA-RAG work. It attributes AuthorityBench to "Yao, Zhang & Bi, CAS State Key Lab of AI Safety."
What needs verifying.
- Does
arXiv:2603.25092resolve, and is it the AuthorityBench paper? If not, find the correct identifier. - Confirm venue (EMNLP 2025? arXiv preprint 2026? both?), authors, and affiliation.
- Confirm the quote: "Incorporating webpage text generally degrades LLM judgment under all settings, indicating authority is not equivalent to textual style, fluency, or narrative richness."
- Confirm the quantitative claim: Qwen3-32B, Spearman ρ 75.28% on the DomainAuth setting.
- Confirm RA-RAG (aclanthology 2025.emnlp-main.1738) is real and does "estimate reliability separately and fuse via weighted majority voting."
Why it matters. This is the load-bearing "reliability wants to be judged blind" result and the machine half of the whole cross-time bridge (observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model). A mis-stamped or fabricated citation would undercut the bridge's strongest claim. A 2026-arXiv-id-under-a-2025-label is exactly the smell worth chasing before this goes budding.
Candidate next move. Fetch the arXiv id directly; if dead, search arXiv/ACL/Semantic Scholar for "AuthorityBench" + "DomainAuth" + the author names; cross-check the ρ figure against the paper's results tables.
Resolution (2026-07-12, cross-model audit, auditor claude-fable-5)
All five items verified against the primary PDF (arXiv:2603.25092, sha256 8bcbef72…, 18 pp.):
- Resolves.
arXiv:2603.25092v1 [cs.IR], submitted 26 Mar 2026, is "AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation." - Venue/authors/affiliation confirmed. arXiv preprint, 2026 — the capture's "2025 research" label was the error, not the id. Authors: Zhihui Yao, Hengran Zhang, Keping Bi; State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (+ UCAS). No EMNLP venue claimed for AuthorityBench itself.
- Quote confirmed verbatim in the paper's §1 ("Incorporating webpage text generally degrades LLM judgment under all settings, indicating authority is not equivalent to textual style, fluency, or narrative richness"). The abstract's phrasing is "consistently degrades judgment performance."
- Figure confirmed. Table 1 (DomainAuth, fine-grained 10-level): Qwen3-32B, w/o webpage text, PairJudge-PointScore → Spearman ρ = 75.28% (also stated in §4.3 prose as "its best score of 0.7528 on DomainAuth"). Caveat now recorded in the claim note: the degradation-from-text holds for ListJudge/PairJudge; PointJudge and hard pairs improve with text (§4.2).
- RA-RAG confirmed real. aclanthology.org/2025.emnlp-main.1738 = Hwang et al., "Retrieval-Augmented Generation with Estimation of Source Reliability," EMNLP 2025, pp. 34279–34303 — estimates source reliability by cross-checking across sources and aggregates via weighted majority voting (WMV). (AuthorityBench's own bibliography cites it with slightly different pages, 34267–34291 — a defect in their bib, not ours.)
Claim note updated accordingly; [unverified-quant] flag cleared with an appended audit_status entry.