ForecastBench (2025) found expert human forecasters significantly outperformed the best-scoring LLM on live, unresolved forecasting questions
ForecastBench (Karger, Bastani, Chen, Jacobs, Halawi, Zhang & Tetlock, ICLR 2025) is a dynamic benchmark of roughly 1,000 live, unresolved forecasting questions, designed specifically to dodge the data-leakage problem that lets a static benchmark's answers leak into an LLM's training data. A 200-question subset scored LLMs against expert human forecasters and the general public on the same live questions. The headline result: "expert forecasters outperform the top-performing LLM (p-value < 0.001)" — a statistically decisive gap, despite LLMs posting superhuman results on many other reasoning and knowledge benchmarks released around the same period.
Tetlock is a co-author. Decades earlier, his own Expert Political Judgment tournaments produced the field's most-cited finding that professional political forecasters barely beat chance — a benchmark a different population of forecasters, intelligence analysts studied by Mandel & Barnes (2014), beat by a wide margin. ForecastBench puts Tetlock on the other side of a structurally similar measurement: instead of asking whether humans beat chance, it asks whether the newest candidate forecaster — the LLM — beats the humans. As of this benchmark, it does not.
Source
“expert forecasters outperform the top-performing LLM (p-value < 0.001)”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-08-12-hop-underconfidence-forecastbench.md, 2026-08-12 · raw markdown