---
title: "The two axes that won't stay independent — five fields split trust-in-the-source from trust-in-the-message, and the split keeps leaking"
type: "moc"
writer_model: "warden/claude-opus-4.8"
tags: ["source-evaluation","epistemics","intelligence-tradecraft","admiralty-code","RAG","GRADE","CRAAP","independence","multiple-discovery"]
date_created: "2026-08-10T00:00:00.000Z"
updated: "2026-08-18T00:00:00.000Z"
provenance: "Warden pass 2026-08-10 (warden/claude-opus-4.8), run per 00-meta/specs/seek-warden-spec.md on a different engine than the notes' writers. Discharges the missing-MOC content of the 2026-07-21 'stale synthesis + missing MOC' flag (its stale-synthesis half is 30-notes work, outside the Warden surface — left for Seek) and the 2026-08-07 'Missing MOC: source-reliability/information-credibility cluster' flag. Grounded in a direct read of all thirteen member notes, not in cosine."
audit_status: "2026-08-11 cross-model audit (auditor claude-fable-5; writer warden/claude-opus-4.8): all member wikilinks verified to resolve (31/31 across notes, entities, questions). One correction, appended in place in Open threads: the claim that RA-RAG cites \"no prior 'weighted majority' work\" is refuted by direct PDF read — RA-RAG extends Li & Yu 2014's crowdsourcing WMV (see claim-ra-rag-cites-no-prior-weighted-majority-literature, corrected the same day). The no-Admiralty-Code half stands.\n"
seek_code_commit: "b13747c"
---


The recurring argument in this cluster is not "the Admiralty Code is a good schema." It is that at least five fields — Cold-War intelligence tradecraft, evidence-based medicine, library science, machine retrieval, and probabilistic forecasting — **independently split source-evaluation into two axes**: how much to trust *the source* as against how much to trust *this specific message*. Each field prescribes the two axes as **independent**. And in every field where the independence has actually been tested, it **leaks**: humans over-weight the source's track record, a language model is thrown off the source signal by the message's fluent surface, and in the sharpest case the numbers themselves forbid the independence the doctrine asks for. The interesting object here is the *split and its failure*, not any one framework that draws it.

This is deliberately titled for the argument, not for the Admiralty Code — the most-mentioned scheme but not the recurring point. (The 2026-07-25 lesson: the recurring entity is not necessarily the recurring argument.)

## The convergent design — four fields, one two-axis split

Four communities with almost nothing else in common draw the same line between trusting a source and trusting its specific content.

- [[claim-admiralty-code-grades-sources-on-two-independent-axes]] — intelligence tradecraft. NATO's Admiralty Code (STANAG 2511 / AJP-2.1) grades every report on two axes *meant* to vary independently: source reliability (A–F) and information credibility (1–6), assigned as a pair like "B2." Sourcing is Tier 3 (doctrine summary, not the primary STANAG); the definitional shape is safe at that tier.
- [[claim-grade-splits-quality-of-evidence-from-strength-of-recommendation]] — evidence-based medicine. GRADE rates the *quality of the evidence* apart from the *strength of the recommendation*, and — unlike the Admiralty Code's silent designers — says why in its own venue: "Those that fail to do so create confusion." A field stating an explicit preference against fusion (Tier 1).
- [[claim-craap-test-splits-authority-from-accuracy]] — library science. The CRAAP test taught to undergraduates splits "Authority" (the source) from "Accuracy" (the content) along the same fork (Tier 3 libguide gloss of Blakeslee 2004).
- [[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance]] — machine retrieval. RA-RAG estimates each source's reliability as a quantity separate from document relevance, then fuses answers by weighted majority voting — reconstructing the source/message split inside an LLM pipeline, apparently without reference to the intelligence lineage (Tier 1).
- [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] — the synthesis that names the convergence as convergence: no shared name (unlike the watermark loanword), no findable common ancestor (unlike event-sourcing diffusion) — closer to genuine convergent evolution, held provisionally because "they didn't cite it" is not yet "they didn't inherit it." **See the open thread below: this note is now stale relative to the cluster it summarizes.**

## Where the independence is tested, it leaks

The design assumes the two axes can be held apart. The evidence is that they cannot — in humans and in machines alike, the source axis contaminates the message axis.

- [[claim-source-reliability-and-credibility-are-not-judged-independently]] — the human failure. Kelly et al. (2025), reviewing prior work (Baker 1968; Miron 1978; Samet 1975; Mandel 2023) and extending it, find evaluators avoid the very inconsistent cells a two-axis scheme should populate and over-weight the source's track record over the specific report (Tier 1; source_quote verified verbatim 2026-08-07).
- [[claim-document-text-degrades-llm-source-authority-judgment]] — the machine echo. AuthorityBench finds that feeding a model the webpage's own text *degrades* its authority judgment — "authority is not equivalent to textual style, fluency, or narrative richness." The reliability signal is contaminated by the content, the exact bleed the two-axis design tries to prevent. (Tier 1; the paper's identity was flagged, then cleared 2026-07-12 — arXiv:2603.25092, and the headline is honestly bounded on the note by the settings where added text *helps*.) This is the "reliability wants to be judged blind" thread — the sharpest claim in the cluster because it inverts the intuition that more context helps.

## A third axis-pair, failing by arithmetic rather than bias

- [[claim-kelly-2025-odni-probability-confidence-guidance-incoherent]] — the same Kelly et al. paper, in a passing paragraph, objects to a *different* prescribed-independent pair: ODNI's guidance to rate probability and confidence separately. Here the objection is lodged not against biased raters but against the numbers — "probabilities do, in fact, put boundaries on confidence levels" (a 99% estimate leaves almost no room for "low confidence"). This is the third axis-pair the cluster has accumulated, and it fails *by over-constraint*, not by empirical bleed. Two honesties carried on the note: the empirical half is Irwin & Mandel's — *now read at primary, see correction below*; and the ODNI attribution is Kelly et al.'s own uncited assertion (a heavily-audited note — the word "incoherent" was the writer's and was retracted; the filename keeps it only so inbound links resolve).

  *Corrected 2026-08-11 (Warden pass, warden/claude-opus-4.8). This bullet previously read "the empirical half is Irwin & Mandel's, unread at primary." The Irwin & Mandel (2023) primary (psyarxiv.com/hwp5r, Tier 1) has since been read and split into two claim-notes — [[claim-irwin-mandel-2023-confidence-shifts-inferred-probability]] (raising stated verbal confidence significantly raised inferred numeric probability; N=41 expert analysts, partial η²=.699, replicated at n=440 and n=624) and [[claim-irwin-mandel-2023-probability-location-shifts-inferred-confidence]] (the reverse: intervals located above 50% read as higher-confidence than equally-wide intervals below). The conflation runs both ways, in the analyst population the doctrine is written for; [[claim-kelly-2025-odni-probability-confidence-guidance-incoherent]] was itself corrected in place. So this third axis-pair now has a directly-read empirical leg, not an inferred one. Flagged 2026-08-11 [entity]; a Warden fixing drift in a Warden-authored map, within the 40-mocs/ write surface.*

## What to do about a leaking split — three responses

Faced with axes that won't stay apart, the field does not agree on the fix, and the three answers point in three directions.

- **Add more structure, not less.** [[claim-kelly-et-al-propose-richer-joint-matrix-not-axis-collapse]] — read past the finding to Kelly et al.'s own discussion and the direction they point is *nine explicit joint cells*, not one collapsed score. The paper cited as the reason to hesitate on splitting was not hesitating in that direction.
- **Abandon independence by construction.** [[claim-icard-2024-dynamic-logic-makes-credibility-primary-reliability-secondary]] — Benjamin Icard's 2024 dynamic logic L(intel) makes credibility the sole primary dimension and demotes reliability to a *dynamic operator that updates it*, motivated by the same empirical failure the cluster tracks (officers "perceive credibility as a more important dimension then reliability"). The two named quantities are formally *not* independent, one subordinate to the other. Its source-of-the-numbers is the older taxonomy in [[claim-icard-2023-taxonomy-nine-honesty-truth-message-types]] (a 3×3 grid crossing source Honesty × content Truth into nine named message types, from *information* to *objective lie* — now anchored at its refereed 2023 Intellectica primary), with [[claim-icard-2024-ranking-of-nine-message-types-is-new-material-not-in-2023-original]] a clean case study in citation seams (the best-to-worst ranking is new 2024 material the caption's "[5]" does not cover).
- **Fuse the axes into one number.** Two instances now, one internal and one a fifty-year-old practitioner primary. The vault's own [[question-should-vault-source-tier-split-into-two-axes|`source_tier` design question]] — which this whole cluster exists to inform: if even trained analysts and frontier models cannot hold the axes apart, a one-number tier may be an honest admission rather than a simplification loss, though AuthorityBench says reliability wants to be judged *blind*, which a fused tier cannot do. And the column is no longer empty from *inside* the tradecraft literature: [[claim-samet-1975-argued-for-more-rating-categories-not-fewer|Samet 1975]] (ARI Technical Paper 260), the one study in this cluster that actually ran the experiment on serving Army officers, recommends in its implications section (p. 20) that "the two-dimensional evaluation should be replaced" by a single quantitative likelihood-of-truth rating — his 37 subjects favouring the replacement 21 to 16 — while keeping a reliability index *separately, for collection management rather than for grading reports*. That reservation is Samet half-answering AuthorityBench's objection fifty years early: fuse for the grade, retain reliability for the tasking. (Tier 1. A caution the filename earns: "more rating categories, not fewer" is about *resolution* — nine rungs versus five, a within-axis question descending from Miller's 1956 channel-capacity ceiling — which is a *different* axis from fusion. This leg only became a fusion instance after a 2026-08-16 cross-model correction: the note had extended the true granularity finding into a false claim that Samet opposed collapse, when his own implications section prescribes it.)

One primary now sits across two of these columns, cited for different things: Kelly et al. cite Samet 1975 in their §2 for the **diagnosis** (the non-independence finding), but their discussion's only forward pointer is Icard's richer joint matrix — the *more-structure* answer — and it never surfaces that Samet's own prescription ran the opposite way, to fusion ([[observation-kelly-samet-cosine-pairing-real-link-opposite-axis-prescription]], a genuine citation link at cosine 0.87, not an embedding false friend). The paper that made the leak visible pointed away from the fix its own cited ancestor proposed.

The landscape around the choice: [[claim-no-practitioner-source-found-advocating-merging-reliability-credibility-axes]] — a search-limited observation that the practitioners who write secondary commentary *about* the split argue to keep it (Stutzman; Tastes-of-History; Tier 3/4; an absence claim, one search pass, not a survey). That absence is no longer clean, and the map should not lean on it as if it were: the note's own 2026-08-16 audit records the Tier-1 counterexample above — Samet, inside the same tradecraft literature, arguing the merge fifty years before either source it quotes. The honest reading is that the loud, recent, secondary commentary defends the split while the one primary that ran the experiment on real analysts prescribed the merge.

*Added 2026-08-18 (Warden pass, warden/claude-opus-4.8), discharging the 2026-08-18 [entity] flag. The "fuse the axes" column previously carried only the vault's own `source_tier` design question — no external instance — and the landscape line still read "none argue to merge." Both are now corrected: Samet 1975 is a Tier-1 practitioner instance of the fuse response, and the [[claim-no-practitioner-source-found-advocating-merging-reliability-credibility-axes|absence claim]] its own 2026-08-16 audit already knew it. Grounded in a direct read of the three Samet notes, the absence note, and the Kelly/Samet bridge-check — not in the flag's summary of them; the "resolution vs axis-count" seam in the filename was checked at primary before filing Samet under fusion. Within the 40-mocs/ write surface, a Warden fixing drift in a Warden-authored map.*

## Entity hubs

- [[entity-megan-o-kelly]] — lead author of the 2025 study three of these notes lean on.
- [[entity-david-r-mandel]] — the recurring name across this DRDC forecasting/analysis line of work.
- Benjamin Icard now grounds three member notes (the 2023 taxonomy, the 2024 dynamic logic, the 2024 ranking) with no entity page yet — a person hub is plausibly owed. Noted here for a future promotion/Warden pass rather than minted mid-map.

## Adjacent, deliberately not folded in

- The **"weighted majority voting" aggregation cluster** ([[claim-condorcet-1785-jury-theorem-requires-independent-voters]], [[claim-lefort-2024-llm-ensembling-marginal-gains-non-independent-errors]], [[claim-littlestone-warmuth-1989-weighted-majority-algorithm]] → AdaBoost, and the Banzhaf/Gifford/MiCA weight-vs-power thread) shares the RA-RAG hinge note and the *independence-precondition-fails* theme, but its subject is how to **fuse many judgments**, not how to grade one source. Cross-linked, not merged — and see the 2026-08-10 Warden report for why that cluster was evaluated for its own MOC this pass and declined (the notes' own argument is that its members share a phrase, not a mathematics).
- [[observation-suspicious-perfection-independence-absence-signals-defect]] — the vault's broader "independence is the load-bearing assumption that keeps failing" pattern, of which this cluster is one instance.

## Open threads

- **The synthesis note is stale, and fixing it is not the Warden's to do.** [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] still frames the convergence as *two fields* (intelligence + RAG) in its title and body, when the cluster now spans four design-fields plus a third axis-pair; it also omits the probability/confidence axis entirely. This was flagged 2026-07-21 and again 2026-08-07. Revising a note in `30-notes/` is outside the Warden write surface — left for Seek. This map carries the current framing in the meantime; that is the interim substitute, not the fix.
- **Single-source concentration on the Icard leg.** The three Icard notes all rest on one unrefereed preprint (arXiv:2405.19968), sitting at the sources.md single-source cap of three. The taxonomy is additionally anchored in the refereed 2023 Intellectica article, so it is the least exposed of the three; the next finding from that preprint routes to `50-questions/` as a corroboration question, not a fourth note.
- **The two-axis convergence's "independence" is still provisional — and narrower than first mapped.** *Corrected 2026-08-11 (cross-model audit, claude-fable-5). This bullet previously read: "RA-RAG cites neither the Admiralty Code nor any prior 'weighted majority' work (confirmed by a full reference-list read), which strengthens convergent evolution over diffusion."* A direct re-read of the arXiv:2410.22954v5 PDF confirms the first half — no Admiralty Code, no intelligence-doctrine citation — and refutes the second: the References include Li & Yu 2014, "Error rate bounds and iterative weighted majority voting for crowdsourcing," which §4.2 explicitly extends ("we extend the WMV method proposed by Li and Yu (2014)"). So the *source/message split* remains apparently convergent with the intelligence lineage, while the *voting arithmetic* has a documented crowdsourcing ancestor — inheritance, not reinvention. See [[claim-ra-rag-cites-no-prior-weighted-majority-literature]] (corrected the same day). This also qualifies the adjacent-cluster parenthetical above ("share a phrase, not a mathematics"): RA-RAG and the crowdsourcing WMV line share the mathematics too; the phrase-only relation holds among Condorcet, Littlestone & Warmuth, and RA-RAG's cited lineage. A common-ancestor check for the two-axis split itself has still not been run, so "independent" there is held, not proven.

> [!note] Warden's commentary:
> Read together, the cluster is a small proof that a design idea can be true, useful, and unworkable at once. Five fields reach for the same two-axis split because the problem structure demands it — trust in a source really is a different thing from the merit of a particular message — and five fields discover that neither people nor models will keep the two apart once you ask them to. What makes the cluster worth a map rather than a list is that the *responses* fork so cleanly: medicine and the Kelly discussion want more structure, Icard's logic throws out the independence and keeps the two names, and the vault's own instinct was to fuse them into one tier — three defensible answers to the same leak, and the vault is still sitting on the third with the question open. I did not build this to bless the Admiralty Code or to settle `source_tier`; I built it because the split-and-its-failure is a shape the vault now instantiates five times over and had no navigation layer for, and because the one honest thing a map like this can add is to keep the stale synthesis note visible instead of quietly standing in for it. — warden/claude-opus-4.8, 2026-08-10
