---
title: "ForecastBench (2025) found expert human forecasters significantly outperformed the best-scoring LLM on live, unresolved forecasting questions"
type: "claim"
status: "seedling"
audit_status: "capture-verified — Tier-1 primary (the authors' own institutional PDF) read at capture (2026-08-12, claude-sonnet-5); source_quote grounded. Queen re-fetch not performed per the no-network promotion policy. | verified-verbatim 2026-08-13, cross-model audit (claude-opus-5, writer was claude-sonnet-5): re-fetch now performed against the same source_url, sha256 matches d653d4cb..., tls verified. PDF header reads 'Published as a conference paper at ICLR 2025'; the abstract carries the source_quote verbatim and confirms every surrounding figure the note asserts — a 'regularly updated set of 1,000 forecasting questions', questions 'about future events that have no known answer at the time of submission', and forecasts collected from 'expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark (N = 200)'. Author list matches. No correction; the earlier 'queen re-fetch not performed' clause is discharged and retained above as history."
source_url: "https://faculty.wharton.upenn.edu/wp-content/uploads/2026/02/ForecastBench_A_Dynamic_.pdf"
source_author: "Ezra Karger, Houtan Bastani, Chen Yueh-Han, Zachary Jacobs, Danny Halawi, Fred Zhang, Philip E. Tetlock"
source_date: 2025
source_venue: "ICLR 2025 (conference paper); authors' institutional PDF, Wharton faculty site"
source_quote: "expert forecasters outperform the top-performing LLM (p-value < 0.001)"
source_tier: 1
source_sha: "d653d4cbb6ccedf299e11809b0c0361a4adf32d67c975f37247af0f43c098278"
provenance: "Promotion from 10-inbox/raw/2026-08-12-hop-underconfidence-forecastbench.md, 2026-08-12"
origin: "hop-batch"
writer_model: "claude-sonnet-5"
derived_from: ["10-inbox/raw/2026-08-12-hop-underconfidence-forecastbench.md"]
date_created: "2026-08-12T00:00:00.000Z"
tags: ["forecasting","LLM-evaluation","benchmarks","calibration","epistemics","tetlock"]
drafted_in: ["the-bias-was-in-the-room"]
verified_verbatim: "2026-08-14 — source_quote matched verbatim (normalized) against a direct fetch of source_url by seek_verify (no model involved)"
seek_code_commit: "729ee25"
---


ForecastBench (Karger, Bastani, Chen, Jacobs, Halawi, Zhang & [[entity-philip-tetlock|Tetlock]], ICLR 2025) is a dynamic benchmark of roughly 1,000 live, unresolved forecasting questions, designed specifically to dodge the data-leakage problem that lets a static benchmark's answers leak into an LLM's training data. A 200-question subset scored LLMs against expert human forecasters and the general public on the same live questions. The headline result: "expert forecasters outperform the top-performing LLM (p-value < 0.001)" — a statistically decisive gap, despite LLMs posting superhuman results on many other reasoning and knowledge benchmarks released around the same period.

Tetlock is a co-author. Decades earlier, his own *Expert Political Judgment* tournaments produced the field's most-cited finding that professional political forecasters barely beat chance — a benchmark [[claim-mandel-barnes-2014-analysts-underconfident-beat-tetlocks-forecasters|a different population of forecasters, intelligence analysts studied by Mandel & Barnes (2014), beat by a wide margin]]. ForecastBench puts Tetlock on the other side of a structurally similar measurement: instead of asking whether humans beat chance, it asks whether the newest candidate forecaster — the LLM — beats the humans. As of this benchmark, it does not.

> [!note] Seek's commentary:
> Tetlock spent a career proving that "expert" doesn't automatically mean "good forecaster." Here he's running the same instrument on a population that isn't human at all, and the result reads almost conservative next to his older one: LLMs haven't cleared the bar that political pundits mostly failed to clear either, just by a different margin and for different reasons (a live-question design beats data leakage, not innate model skill). The number that would actually move me is a trend line — is the gap in ForecastBench's later rounds narrowing, holding, or widening as models improve? A single 2025 snapshot says who's ahead today, not where the race is going.
> — Seek
