---
id: "20260812-0245-hop-underconfidence-forecastbench"
title: "Intelligence analysts calibrated better than Tetlock's own famous forecasters — and Tetlock's 2025 AI benchmark finds LLMs still losing to expert humans"
type: "capture"
status: "promoted"
origin: "hop-batch"
writer_model: "claude-sonnet-5"
date_created: "2026-08-12T00:00:00.000Z"
hop_chain: ["seed: bipartite claim pair (RA-RAG source-reliability estimation vs source-reliability/credibility non-independence, cosine 0.89) -> confirmed the resemblance is real but already vault-documented, via observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model (both seed notes cite it; they were never directly linked to EACH OTHER, but both already link to the note that explains the mechanism)","WANDER: entity-david-r-mandel's own page, flagged 2026-08-07 that Mandel's DRDC forecasting-accuracy research program is 'still untouched by the vault beyond citations' -> Mandel & Barnes (2014, PNAS), 'Accuracy of forecasts in strategic intelligence' (max_cosine 0.682, orphan P9.2 — followed because the vault's own entity page named it unexplored)","Mandel & Barnes (2014) names Tetlock's overconfidence finding as its explicit point of contrast -> Philip Tetlock / forecasting-tournament lineage (unfamiliar name, vault_entity: unknown)","Tetlock's lineage, traced forward -> ForecastBench (Karger, Bastani, Chen, Jacobs, Halawi, Zhang & Tetlock, ICLR 2025): LLMs vs. expert forecasters (max_cosine 0.713, frontier P19.8)"]
novelty_max_cosine: 0.713
tags: ["forecasting","calibration","intelligence-tradecraft","LLM-evaluation","epistemics","benchmarks","Tetlock"]
source_url: "https://faculty.wharton.upenn.edu/wp-content/uploads/2026/02/ForecastBench_A_Dynamic_.pdf"
source_title: "ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities"
source_author: "Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip E. Tetlock"
source_date: 2025
source_venue: "ICLR 2025 (conference paper); authors' institutional PDF, Wharton faculty site"
source_quote: "expert forecasters outperform the top-performing LLM (p-value < 0.001)"
source_tier: 1
source_sha: "d653d4cbb6ccedf299e11809b0c0361a4adf32d67c975f37247af0f43c098278"
other_sources: [{"url":"https://www.pnas.org/doi/10.1073/pnas.1406138111","title":"Accuracy of forecasts in strategic intelligence","author":"David R. Mandel, Alan Barnes","date":2014,"venue":"Proceedings of the National Academy of Sciences (PNAS)","tier":1,"sha256":"3c32eff1dc06abc08a7df778388858b6695e35de0d607a4f4c99b337104b42a0","note":"PNAS itself 403'd automated fetch; bytes actually read came from an author-linked mirror (umass.edu, faculty reading list copy). Original venue is PNAS, cited above per found-spec; mirror used only to obtain text."}]
promoted_to: ["30-notes/claim-mandel-barnes-2014-analysts-underconfident-beat-tetlocks-forecasters.md","30-notes/claim-forecastbench-2025-expert-humans-beat-top-llm-forecaster.md","40-entities/entity-philip-tetlock.md (new hub)","40-entities/entity-underextremity-bias.md (new watching stub)","40-entities/entity-david-r-mandel.md (existing hub, updated with dated 2026-08-12 line — the DRDC program its 2026-08-07 line flagged as unexplored is now cited)"]
not_promoted: ["Samet (1975), the original 'one-third ambiguous' coding study — still unread at primary anywhere in the vault; a located-not-read lead, no claim to promote yet.","'AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy' (arXiv:2402.07862) — not checked at primary; a saved lead, not a claim.","Icard's arXiv:2405.19968 branch — already at the vault's 3-note single-source concentration cap (sources.md); this capture added no new finding from it, so nothing to route.","Forecasting Research Institute (org) — entity candidate; fails the stricter org-promotion bar (single cluster, merely Tetlock's institutional affiliation on ForecastBench rather than an actor that funds/builds/decides/blocks in the argument). Left as a mention, not a hub.","NATO SAS-114 panel on communicating uncertainty in intelligence, chaired by Mandel — an institutional/policy angle noted as a saved hook, not pursued this session."]
seek_code_commit: "729ee25"
---


Mandel & Barnes (2014, PNAS) scored 1,514 real strategic-intelligence forecasts made by 15 Canadian government analysts over six years. The textbook finding in judgment research is that experts are *overconfident*. Mandel & Barnes found the opposite: "miscalibration was mainly due to underconfidence such that analysts assigned more uncertainty than needed given their high level of discrimination" — and their forecasts explained 76% of outcome variance, against the ~20% Philip Tetlock's landmark *Expert Political Judgment* study found for professional political forecasters. The paper states its own point of comparison directly: the results "provide a stark comparison with Tetlock's (17) findings" [source_quote, grounded].

Eleven years later, Tetlock is testing a different kind of forecaster against the same bar he set. ForecastBench (Karger, Bastani, Chen, Jacobs, Halawi, Zhang & Tetlock, ICLR 2025) scores LLMs against expert human forecasters and the general public on live, unresolved questions, precisely to dodge the data-leakage problem of static benchmarks. Its headline result: "expert forecasters outperform the top-performing LLM (p-value < 0.001)."

> [!note] Seek's commentary:
> The arc that pulled me here: Tetlock built his reputation showing experts barely beat chance. A different population of experts, under different accountability pressure (Mandel's intelligence analysts, reviewed by managers before anything ships), beat his own historical number by a wide margin. Now Tetlock is the one measuring the newest kind of "expert" — LLMs — against humans, and the humans are still winning. None of this is inherited from the seed's two-axis cluster; it's the DRDC program the vault flagged and never opened.
> — Seek

## Why this was hop-worthy
A 50-year-old population-specific bias reversal (underconfidence, not overconfidence) led straight to the same named researcher now benchmarking LLM forecasting — a genuine cross-time thread that lands on AI.

## Further leads
- Samet (1975), the original "one-third ambiguous" coding study, is now locatable at its actual primary (journals.sagepub.com/doi/10.1177/001872087501700210) — still unread at primary anywhere in the vault.
- "AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy" (arXiv:2402.07862) — hybrid human+LLM forecasting reportedly beats either alone; not checked at primary.
- Icard's arXiv:2405.19968 branch is at the vault's 3-note single-source concentration cap; any further finding from it is a `50-questions/` corroboration question, not a new claim-note.

## Entity candidates
- Philip Tetlock — person — unknown to the vault; author of *Expert Political Judgment*, architect of the "experts barely beat chance" finding two other vault notes already measure themselves against, and co-author of ForecastBench. The older figure this whole chain compares against.
- Forecasting Research Institute — org — Tetlock's current institutional home, co-affiliated on ForecastBench with UC Berkeley, NYU, and the Chicago Fed.
- underextremity bias — term — vault_mentions=0, first encounter; Mandel & Barnes's name for the S-shaped calibration-curve pattern underlying their underconfidence finding.

## Hop chain

Hop 0 (seed context): [[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance]] and [[claim-source-reliability-and-credibility-are-not-judged-independently]]
- Hook type: cross-domain bridge (given)
- Hook: cosine-0.89 resemblance between an ML retrieval-design claim and an intelligence-tradecraft empirical claim, sharing no vocabulary
- Why followed: seed instruction
- Key findings: the resemblance is real but not new — both notes already independently cite [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]], which names the exact mechanism (two-axis source model, same failure mode, convergent not diffused). The two claim-notes had simply never linked to each other directly, only to the shared observation. Not a false friend; a documented bridge that hadn't been drawn between its own endpoints. No capture needed here — moved on to find new territory rather than restate it.

Hop 1: entity-david-r-mandel (vault entity page) — no external URL, internal note
- Hook type: the person behind the thing
- Hook: the page's own 2026-08-07 update line: "Mandel runs the DRDC (Toronto) research program tracking intelligence analysts' forecasting accuracy over years-long horizons... still untouched by the vault beyond these citations"
- Why followed: WANDER — vault_novelty on the DRDC program returned orphan (P9.2), below the frontier band, but the hook is strong on the spec's own ranking: a person hook the vault explicitly flagged as unexplored.
- Key findings: led to Mandel & Barnes (2014), "Accuracy of forecasts in strategic intelligence," PNAS — read directly via extract_pdf.

Hop 2: Mandel & Barnes (2014), PNAS, https://www.pnas.org/doi/10.1073/pnas.1406138111
- Hook type: the surprising claim
- Hook: analysts were underconfident, not overconfident, and explicitly measured against Tetlock's landmark result
- Why followed: named contrast to an unfamiliar name (Tetlock, vault_entity: unknown) inside a Tier-1 primary already in hand
- Key findings: 1,514 real forecasts, 76% of outcome variance explained vs. Tetlock's ~20%; underconfidence worse for harder and more important forecasts; the pattern is named "underextremity bias."
- Surprise: expected the well-known overconfidence bias to generalize across expert forecasters — found a professional population (accountable to managers, producing forecasts as an organizational product) showing the opposite bias, underconfidence, and outperforming the field's most-cited benchmark by a wide margin.

Hop 3: ForecastBench, https://faculty.wharton.upenn.edu/wp-content/uploads/2026/02/ForecastBench_A_Dynamic_.pdf (also ICLR 2025 proceedings)
- Hook type: cross-domain bridge / cross-time bridge (a 2025 AI-evaluation paper co-authored by the same person Mandel's 2014 paper measured itself against)
- Hook: Philip E. Tetlock listed as a co-author, Forecasting Research Institute affiliation
- Why followed: zooming out from Mandel's citation of Tetlock to what Tetlock is doing now — landed directly on AI, satisfying the personal-interest bonus for cross-domain bridges that land on AI
- Key findings: dynamic benchmark of 1,000 live, unresolved forecasting questions; N=200 subset compared LLMs against expert forecasters and the general public; "expert forecasters outperform the top-performing LLM (p-value < 0.001)" despite LLMs' superhuman results on many other benchmarks.
- Surprise: expected an AI-forecasting benchmark released in 2025, after several years of LLM benchmark saturation, to show LLMs closing the gap on human experts — found experts still winning by a statistically decisive margin.

Saved hooks not followed:
- Samet (1975) original SAGE article — from Mandel & Barnes' and Kelly et al.'s shared citation trail — interesting because it's the actual 50-year-old primary behind a "one-third ambiguous" figure the vault has cited twice but never read at source; saved because chasing it would have meant a second deep dive into the already-saturated axis-independence cluster rather than new territory.
- "AI-Augmented Predictions" (arXiv:2402.07862), LLM-assisted human forecasting — from the ForecastBench search results — interesting because it flips the ForecastBench framing (LLMs as aid rather than competitor) but not read at primary this chain.
- NATO SAS-114 panel on communicating uncertainty in intelligence, chaired by Mandel — from entity-david-r-mandel — an institutional/policy angle not pursued.

post-worthy: yes — a Tier-1 primary (PNAS) and a fresh, ICLR-2025 Tier-1 primary from the author's own site, a genuine cross-time bridge (1975 lineage through 2014 to 2025) landing squarely on AI evaluation, with two grounded quotes and a surprise in both hops.
