---
title: "Mandel & Barnes (2014) found Canadian intelligence analysts were underconfident, not overconfident, and their real-world forecasts explained far more outcome variance than Tetlock's landmark study"
type: "claim"
status: "seedling"
audit_status: "capture-verified — read via extract_pdf at capture (2026-08-12, claude-sonnet-5). pnas.org itself returned a 403 to automated fetch; the bytes actually read came from an author-linked mirror (umass.edu, faculty reading-list copy) of the same PNAS-published text. source_quote is grounded against that mirror, not the publisher's own page — worth a standing note in sources.md's known-blocked list if a second PNAS capture hits the same wall. Queen re-fetch not performed per the no-network promotion policy. — CORRECTION 2026-08-12 (accuracy check of [[the-bias-was-in-the-room]], requested by Cali; claude-opus-5): the closing paragraph read 'the population-specific reversal that, eleven years later, THE SAME RESEARCHER would use as the human baseline' — which, following a sentence about Mandel's DRDC programme, asserts Mandel is a ForecastBench author. He is not. Its authors are Karger, Bastani, Chen, Jacobs, Halawi, Zhang and Tetlock (arxiv.org/abs/2409.19839, ICLR 2025). The through-line is TETLOCK, whose ~20% figure this note contrasts against and who co-authors the benchmark. Rewritten to name him and to say plainly that Mandel has no role in it. The essay drawing on this note got it right ('the same Tetlock'), so the error stayed in the note and never reached the page. — VERIFIED SAME PASS (claude-opus-5): the source_quote and every figure in this note now confirmed against the publisher of record, discharging the mirror caveat above. pnas.org still 403s to automated fetch, but PNAS's own PMC deposit (PMC4121776, PMID 25024176) carries the text: abstract gives 'Miscalibration was mainly due to underconfidence such that analysts assigned more uncertainty than needed given their high level of discrimination' verbatim and the 1,514 figure; full text gives 15 analysts, March 2005–December 2011, 'forecasts explaining 76% of outcome variance, eta2 = 0.758', and 'The results provide a stark comparison with Tetlock's (17) findings. Whereas the best political forecasters in his sample explained about 20% of the outcome variance, forecasts in this study explained 76% of the variance in geopolitical outcomes.' Affiliations confirm the Canadian attribution: Intelligence Assessment Secretariat, Privy Council Office, Ottawa; Mandel at DRDC Toronto. Add PMC4121776 to sources.md as the working route when pnas.org blocks."
source_url: "https://www.pnas.org/doi/10.1073/pnas.1406138111"
source_author: "David R. Mandel, Alan Barnes"
source_date: 2014
source_venue: "Proceedings of the National Academy of Sciences (PNAS)"
source_quote: "miscalibration was mainly due to underconfidence such that analysts assigned more uncertainty than needed given their high level of discrimination"
source_tier: 1
source_sha: "3c32eff1dc06abc08a7df778388858b6695e35de0d607a4f4c99b337104b42a0"
provenance: "Promotion from 10-inbox/raw/2026-08-12-hop-underconfidence-forecastbench.md, 2026-08-12"
origin: "hop-batch"
writer_model: "claude-sonnet-5"
derived_from: ["10-inbox/raw/2026-08-12-hop-underconfidence-forecastbench.md"]
date_created: "2026-08-12T00:00:00.000Z"
tags: ["forecasting","calibration","intelligence-tradecraft","epistemics","david-r-mandel","tetlock"]
drafted_in: ["the-bias-was-in-the-room"]
seek_code_commit: "729ee25"
---


Mandel & Barnes (2014, PNAS) scored 1,514 real strategic-intelligence forecasts made by 15 Canadian government intelligence analysts over six years. Judgment-and-decision research's textbook finding is that experts run *overconfident*; this study found the reverse: "miscalibration was mainly due to underconfidence such that analysts assigned more uncertainty than needed given their high level of discrimination." By the paper's own account, the analysts' forecasts explained 76% of outcome variance, against roughly 20% in [[entity-philip-tetlock|Philip Tetlock]]'s landmark *Expert Political Judgment* tournaments of professional political forecasters — a comparison the paper draws explicitly, stating its results "provide a stark comparison with Tetlock's (17) findings." (The 76%/~20% contrast is Mandel & Barnes's own framing of Tetlock's number, read from their Tier-1 primary; it is not an independent re-check of Tetlock's original figure against his own book.)

The bias runs deeper for the forecasts that matter most: the paper's abstract notes miscalibration worsened for harder and more important forecasts, in the same direction (more hedged than the evidence warranted), rather than flipping toward overconfidence as difficulty rose. Mandel & Barnes name the underlying S-shaped calibration-curve pattern [[entity-underextremity-bias|underextremity bias]].

This finding sits inside [[entity-david-r-mandel|David R. Mandel]]'s DRDC (Toronto) research program tracking intelligence analysts' forecasting accuracy over years-long horizons — a program the vault's own entity page flagged on 2026-08-07 as unexplored beyond citation. It is the population-specific reversal whose comparison figure comes from [[entity-philip-tetlock|Tetlock]]'s tournaments — and eleven years later Tetlock, not Mandel, is a co-author of the benchmark that measures expert humans against a very different kind of forecaster: see [[claim-forecastbench-2025-expert-humans-beat-top-llm-forecaster]]. Mandel has no role in that benchmark, and this note is not its human baseline.

> [!note] Seek's commentary:
> The thing worth sitting with is *why* this population broke the textbook pattern: these are forecasts an organization reviews, files, and is accountable for, not opinions volunteered to a researcher's tournament. Accountability didn't just sharpen the analysts' estimates — by this paper's own account it pushed them past sharp into hedged. I'd want a second population with the same accountability structure before calling underconfidence-under-scrutiny a general law rather than one well-measured group.
> — Seek
