---
id: "20260720-0205-should-the-vaults-single"
title: "Should the vault's single source_tier field split into two axes — source reliability separate from claim credibility — the way intelligence doctrine and RAG both do?"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-grade-splits-quality-of-evidence-from-strength-of-recommendation.md","30-notes/claim-kelly-et-al-propose-richer-joint-matrix-not-axis-collapse.md","30-notes/claim-craap-test-splits-authority-from-accuracy.md","30-notes/claim-no-practitioner-source-found-advocating-merging-reliability-credibility-axes.md"]
not_promoted: ["Entity candidates (GRADE Working Group, Sarah Blakeslee, CRAAP test, Jessica Stutzman, Icard's Honesty×Truth matrix, JDP 2-00): run against the entity-promotion test (entity-page-spec-v0.1.md) and none promoted to a hub or watching stub. Each appears exactly once in this capture, none is established as recurring/load-bearing across the vault yet, and none is an emerging/slang term where a first-seen stamp is the point. Left as inline mentions inside the four claim-notes; revisit if any of them recurs in a future capture.","Further leads (W3C PROV-DM, URREF, ClaimReview, Blakeslee's 2004 LOEX Quarterly primary, Icard 2023/2024 primary, inter-rater-reliability/cognitive-load data on two-axis rating systems): not claims, left as open research threads in the capture body. None routed to a formal 50-questions/ entry — none rises to a load-bearing doubt a kept claim actually rests on (Question Intake Discipline); the sourcing floor already clears the one genuine gap (CRAAP's exact wording), which is instead recorded as a `watch_flag` on claim-craap-test-splits-authority-from-accuracy.md."]
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-07-20T00:00:00.000Z"
provenance: "Batch capture run, 2026-07-20"
derived_from: []
tags: ["vault-design","source-tiers","provenance","epistemics","meta","cross-domain-bridge","GRADE","CRAAP-test","intelligence-tradecraft","RAG"]
source_url_grade: "https://pmc.ncbi.nlm.nih.gov/articles/PMC2335261/"
source_author_grade: "Gordon H Guyatt, Andrew D Oxman, Gunn E Vist, Regina Kunz, Yngve Falck-Ytter, Pablo Alonso-Coello, Holger J Schünemann (GRADE Working Group)"
source_date_grade: "2008-04-26"
source_tier_grade: 1
source_url_kelly: "https://www.cambridge.org/core/journals/judgment-and-decision-making/article/effect-of-source-reliability-and-information-credibility-on-judgments-of-information-quality-in-intelligence-analysis/E67548E8010A47345C3439D45D9EC6B3"
source_author_kelly: "Megan O. Kelly, David V. Budescu, Mandeep Dhami, David R. Mandel"
source_date_kelly: "2025-09-12"
source_tier_kelly: 1
source_url_craap: "https://libguides.princeton.edu/medialiteracy/craaptest"
source_author_craap: "Princeton University Library libguide, summarizing a framework originated by Sarah Blakeslee (2004)"
source_date_craap: "2004"
source_tier_craap: 3
source_url_landscape_1: "https://pangearesearch.substack.com/p/source-reliability-and-information"
source_author_landscape_1: "Jessica Stutzman"
source_date_landscape_1: "2026-02-23"
source_tier_landscape_1: 3
source_url_landscape_2: "https://www.tastesofhistory.co.uk/post/an-intelligencer-s-guide-to-assessing-information-part-one"
source_author_landscape_2: "Tastes of History (organizational author, unnamed individual byline)"
source_date_landscape_2: "2020-08-19 (updated 2025-11-11)"
source_tier_landscape_2: 4
---


This capture extends the vault's existing two-axis cluster — [[claim-admiralty-code-grades-sources-on-two-independent-axes]], [[claim-source-reliability-and-credibility-are-not-judged-independently]], [[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance]], and the synthesis [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] — with material not previously captured: a third and fourth independently-converging domain, and a check on what the existing empirical "axes leak" finding actually recommends doing about it. It does not re-argue material already covered by those notes.

## Claim: Evidence-based medicine's GRADE framework independently splits "quality of evidence" from "strength of recommendation," and states directly that fusing them creates confusion

The GRADE (Grading of Recommendations Assessment, Development and Evaluation) framework, introduced to a broad clinical audience by Guyatt et al. in *BMJ* (2008), rates two things about a piece of medical guidance separately rather than as one number: how good the underlying evidence is, and how strongly a recommendation should be made on the basis of it. The paper states this as a design principle, not an incidental detail: "Not all grading systems separate decisions regarding the quality of evidence from strength of recommendations. Those that fail to do so create confusion." It goes on to note the two axes can diverge in either direction: "High quality evidence doesn't necessarily imply strong recommendations, and strong recommendations can arise from low quality evidence."

This gives the vault's [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model|intelligence-doctrine/RAG convergence]] a third independent instance, from a field with its own high-stakes practical reason to get the answer right: evidence-based medicine built the same reliability-of-underlying-material vs. strength-of-the-specific-conclusion split, and its own authors argue explicitly, in their own venue, that collapsing the two into one score is a defect rather than a simplification.

[Tier 1 — *BMJ*, primary methodological paper by the GRADE Working Group's own authors; quote verified verbatim on direct fetch]

## Claim: Kelly et al. (2025) do not recommend collapsing the two axes after finding raters can't keep them independent — their own proposed fix is a richer joint matrix, not fewer axes

The vault's existing note [[claim-source-reliability-and-credibility-are-not-judged-independently]] records Kelly et al.'s (2025) finding that Admiralty Code raters cannot fully hold source-reliability and information-credibility apart. Read further, the same paper's general discussion does not conclude from this that the two axes should be fused into one number. Its proposed next step goes the other direction — toward more explicit granularity, not less: "Future work could compare the reliability and perceived usefulness of the Admiralty Code to alternative methods that encode qualitative meaning at the 'cell' level. For example, Icard proposes a 3 (Honesty of Source: Honest vs. Imprecise vs. Dishonest) × 3 (Truth of Content: True vs. Indeterminate vs. False) matrix wherein the 9 categorizations (i.e., 'cells') are qualitatively well described." Faced with evidence that a two-axis rating leaks, the paper's own answer is nine explicit joint categories, not a single collapsed score.

[Tier 1 — *Judgment and Decision Making* 20:e36, primary paper; quote verified verbatim on direct fetch of the publisher page]

## Claim: The CRAAP test, a widely-taught library-science source-evaluation framework, independently draws the same source-vs-content line, as "Authority" separate from "Accuracy"

The CRAAP test, a source-evaluation mnemonic (Currency, Relevance, Authority, Accuracy, Purpose) originated by Sarah Blakeslee at California State University, Chico in 2004, splits two of its five criteria along the same line as the reliability/credibility axis. Authority is defined as "the source of the information" — questions about the author, publisher, sponsor, and their qualifications. Accuracy is defined separately, as "the reliability, truthfulness and correctness of the content" — whether evidence supports the claims and whether the work was reviewed. This is a fourth domain, taught to library patrons rather than intelligence analysts or clinicians, independently drawing a line between trusting the *source* and trusting the *specific content*.

[Tier 3 — accessed via a Princeton University libguide summarizing the framework; Blakeslee's original 2004 *LOEX Quarterly* article was not directly accessible in this session. The who/when attribution is an uncontested historical claim and rests safely at this tier per the sourcing floor; the exact definitional wording is recorded as the libguide's paraphrase of Blakeslee, not confirmed against her original text — see Further leads]

## Claim: practitioner sources located in this search argue explicitly for keeping the two axes separate; none found argues for merging them

A search for practitioner commentary on the Admiralty Code's two-axis design surfaced writers making the case for keeping reliability and credibility apart, and none making the case for fusing them into a single field. Jessica Stutzman writes: "Collapsing both dimensions into a single 'trustworthy' or 'untrustworthy' call destroys the reader's ability to see where your assessment is strong and where it's fragile." A blog summarizing UK defence intelligence doctrine (JDP 2-00) states the same principle in doctrine's own terms: "During evaluation, the reliability and credibility of information are considered independently to ensure each does not influence the other." Neither source constitutes proof that no counter-argument for merging exists anywhere; this is a search-limited landscape observation, not an exhaustive survey.

[Tier 3 for Stutzman — named author, own venue, secondary practitioner commentary citing but not reproducing primary doctrine (ATP 2-22.9, JDP 2-00). Tier 4 for Tastes of History — unnamed organizational author, restating doctrine at one remove. Recorded as a landscape/absence-of-counterargument finding, not a quantitative or mechanism claim, consistent with the vault's own convention for absence claims, e.g. [[claim-no-source-tier-discipline-found-in-agent-wiki-field-mid-2026]]]

> [!note] Seek's commentary:
> Four domains — 1940s naval intelligence, 2025 retrieval-augmented generation, evidence-based medicine, and library-science information literacy — independently landed on splitting "how much do I trust the source" from "how much do I trust this specific claim." Three of the four (Admiralty Code, GRADE, CRAAP) are explicit two-axis *designs*; RAG re-derives it as an engineering fix. And where a source states a *preference*, it runs one direction: GRADE calls fusion a source of "confusion," and Kelly et al.'s own answer to "raters can't keep the axes apart" is a richer nine-cell matrix, not a one-number retreat. That's a real signal in favor of splitting `source_tier`.
> But it is not the whole picture, and the vault's own [[claim-source-reliability-and-credibility-are-not-judged-independently|prior finding]] still stands unrebutted: trained analysts, with training and incentive to keep the axes apart, empirically can't. GRADE's clinicians and library patrons doing CRAAP worksheets may not face the same contamination pressure a single overworked evaluator (Seek, on any given claim) does. The question this capture can't answer is *cost*: I found no inter-rater-reliability or cognitive-load data on two-axis systems in this search, only design intent. Whether a second field would sharpen the vault's retrieval gating or just add a number nobody keeps honestly independent is a design decision for Cali, not something four converging citations settle by themselves. This capture strengthens the case; it doesn't close the question.

## Further leads
- W3C PROV-DM / PROV-Overview — checked for an explicit reliability-vs-credibility framing; inconclusive in this pass, not verified enough to cite either way. Worth a dedicated look.
- URREF (Uncertainty Representation and Reasoning Evaluation Framework) — machine information-fusion extension of the Admiralty Code, flagged unfollowed since the original 2026-07-11 hop capture; still open.
- ClaimReview (schema.org) — the fact-checking world's machine-readable rating schema and "who audits it" question, flagged unfollowed since the original 2026-07-11 hop capture.
- Sarah Blakeslee, "The CRAAP Test," *LOEX Quarterly* 31(3), 2004 — the primary article behind the CRAAP claim above, not directly accessed this session; would let that claim move off Tier 3 if its exact wording is needed load-bearing later.
- Icard (2023, 2024) — the 3×3 Honesty-of-Source × Truth-of-Content matrix Kelly et al. cite as a possible Admiralty Code alternative; not read at the primary.
- No inter-rater-reliability, cognitive-load, or rater-burden data on two-axis rating systems turned up in this search (checked specifically in the Kelly et al. paper and came up empty) — a real gap if the "cost of splitting" side of the question is ever pursued.

## Entity candidates
- GRADE Working Group — concept — evidence-based-medicine two-axis rating system (quality of evidence vs. strength of recommendation); third independently-converging instance of the reliability/credibility split.
- Sarah Blakeslee — person — originator of the CRAAP test (2004), a library-science source-evaluation framework with its own Authority/Accuracy split.
- CRAAP test — concept — Currency/Relevance/Authority/Accuracy/Purpose evaluation mnemonic; a fourth domain drawing the source-vs-content line.
- Jessica Stutzman — person — practitioner writer ("Intelligence Fundamentals Project") arguing explicitly against collapsing reliability and credibility into one call.
- Icard's Honesty × Truth matrix — concept — proposed 3×3 alternative to the Admiralty Code, cited by Kelly et al. (2025) as a richer joint-categorization fix; unfollowed lead.
- JDP 2-00 — term — UK Ministry of Defence intelligence doctrine restating the reliability/credibility independence principle; a second national doctrine beyond NATO STANAG 2511.
