---
title: "In Irwin & Mandel's (2023) N=41 expert-analyst experiment, raising stated verbal confidence significantly increased analysts' inferred numeric probability, not just their interval's margin of error"
type: "claim"
status: "seedling"
audit_status: "capture-verified | 2026-08-13 cross-model audit (claude-opus-5, writer was claude-sonnet-5): source_quote found verbatim in the accepted manuscript (Results 2.2.2.1, expert sample), together with every statistic the note reports — expert ANOVA on interval midpoints F(1, 39) = 90.62, p < .001, partial eta-squared = .699, all pairwise comparisons significant on Fisher's LSD; expert n = 41, made up of 21 analysts recruited during course time at the Canadian Forces School for Military Intelligence and 20 who participated remotely by Qualtrics link; the varying-degrees-of-confidence task wording confirmed as a single 'likely' assessment with 95%-certainty ranges elicited under low, moderate and high stated confidence; non-expert samples n = 440 (Experiment 1) and N = 624 (Experiment 2); Experiment 2 midpoint ANOVA confidence level F(2, 621) = 38.43 against probability level F(1, 621) = 33.83 with interaction F(2, 621) = 13.47, all p < .001, the paper stating 'confidence level influenced the midpoint probability more strongly than probability level'; and the closing quotation verbatim in the General Discussion. This DISCHARGES the open NO MATCH flag raised against this note by seek_verify in 00-meta/reports/verify-2026-08-12.md, which classed the quote as 'artifact/fabrication class' — the cause was URL drift, not fabrication (see below). POINTER CAVEAT: source_url now 301s to https://osf.io/preprints/psyarxiv/hwp5r, whose primary file is a .docx, so the URL resolves to the preprint record but no longer serves a fetchable PDF, and the Wiley version of record returns HTTP 402 to automated fetch. An independent re-download was therefore not possible and the status is deliberately NOT upgraded to verified-verbatim; source_venue added so the citation reaches the version of record."
source_url: "https://psyarxiv.com/hwp5r"
source_author: "Daniel Irwin, David R. Mandel"
source_date: 2023
source_venue: "PsyArXiv preprint (accepted manuscript, marked 'In press: Risk Analysis'); version of record: Risk Analysis 43(5):943-957, 2023, doi:10.1111/risa.14009"
source_quote: "Had experts treated probability and confidence as independent constructs, we would expect to observe invariant midpoint interpretations, yet midpoint interpretations substantially increased with each increase in confidence level"
source_tier: 1
source_sha: "2848a23bbd421bf10eb93d5daa33781dadf8d515b7abd8a3f23c4c43eee89058"
provenance: "Promotion from 10-inbox/raw/2026-08-11-does-irwin-mandel-2023-actually-show-that-intelligence.md, 2026-08-11"
origin: "batch"
writer_model: "claude-sonnet-5"
derived_from: ["10-inbox/raw/2026-08-11-does-irwin-mandel-2023-actually-show-that-intelligence.md"]
date_created: "2026-08-11T00:00:00.000Z"
tags: ["intelligence-tradecraft","epistemics","forecasting","independence","probability-confidence","david-r-mandel"]
seek_code_commit: "b13747c"
---


[[entity-daniel-irwin|Daniel Irwin]] and [[entity-david-r-mandel|David R. Mandel]] gave a sample of 41 professional Canadian intelligence analysts (21 recruited in person, 20 via a remote link) a fixed hypothetical assessment using the term "likely," then asked for 95%-certainty probability ranges under three stated-confidence conditions — low, moderate, high. If confidence were a genuinely independent quantity expressing only margin of error, the inferred range's *midpoint* should stay constant across conditions while only its *width* varies. Instead: "Had experts treated probability and confidence as independent constructs, we would expect to observe invariant midpoint interpretations, yet midpoint interpretations substantially increased with each increase in confidence level" (F(1, 39) = 90.62, p < .001, partial η² = .699; all pairwise comparisons significant). The pattern replicated at larger scale in non-expert samples (n = 440 and n = 624), and in the second sample confidence level had a *stronger* effect on inferred probability than the stated probability term itself (F[2,621] = 38.43 vs. F[1,621] = 33.83, both p < .001, plus a significant interaction). The paper states the conclusion directly: "Critically, we found that intelligence consumers do not treat confidence and probability as independent constructs."

This is the empirical leg [[claim-kelly-2025-odni-probability-confidence-guidance-incoherent|Kelly et al. (2025) cite]] for their claim that "analysts and nonexperts alike do treat probability and confidence as related constructs," and it resolves [[question-verify-irwin-mandel-2023-probability-confidence-cue]]. It joins the vault's [[moc-two-axes-that-wont-stay-independent|recurring pattern of prescribed-independent axes that leak]] as the sharpest empirical case: the leak runs in the direction of confidence contaminating probability, not the reverse (see [[claim-irwin-mandel-2023-probability-location-shifts-inferred-confidence]] for the reverse direction).

> [!note] Seek's commentary:
> A partial η² of .699 is not a subtle effect — confidence level is explaining most of the variance in what should, by design, be an invariant number. What makes this the cleanest instance in the two-axes cluster is that it isn't an argument about over-constraint (like the Kelly 99%-probability case) or an inference from absence (like the no-citation-lineage finding elsewhere in the vault); it's a direct manipulation with a huge, repeatedly-replicated effect size, in the population — trained analysts — the doctrine is actually written for.
> — Seek
