---
title: "William R. Thompson"
type: "entity"
entity_kind: "person"
status: "hub"
canonical_name: "William R. Thompson"
aliases: ["Thompson"]
first_seen: "2026-09-15T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["multi-armed bandit problem","Thompson sampling","exploration-exploitation tradeoff","Bayesian reinforcement learning","A.A. Fel'dbaum / dual control theory"]
seek_code_commit: "546fa57"
---


American physician and statistician who in 1933 gave the first mathematical formulation of the multi-armed bandit problem, in the context of clinical trials — the origin point of "Thompson sampling," a foundational exploration algorithm still in active use in modern reinforcement learning and online experimentation.

Matters to this vault as the independent, decades-older root of a lineage that a 2026 control-theory survey's own citation structure keeps separate from [[entity-aa-feldbaum|A.A. Fel'dbaum]]'s dual control theory, even though the same survey credits Fel'dbaum as exploration-exploitation's origin point in its abstract: the survey's own multi-armed-bandit section traces its ancestry to Thompson in 1933, three decades before Fel'dbaum, and never names Fel'dbaum within it ([[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]). Mainstream RL's actual exploration algorithms — epsilon-greedy, UCB, Thompson sampling itself — descend from this 1933 statistics/operations-research lineage, not from Fel'dbaum's control-theoretic risk decomposition.

## References
- [[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]
- Capture: 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md
