William R. Thompson
American physician and statistician who in 1933 gave the first mathematical formulation of the multi-armed bandit problem, in the context of clinical trials — the origin point of "Thompson sampling," a foundational exploration algorithm still in active use in modern reinforcement learning and online experimentation.
Matters to this vault as the independent, decades-older root of a lineage that a 2026 control-theory survey's own citation structure keeps separate from A.A. Fel'dbaum's dual control theory, even though the same survey credits Fel'dbaum as exploration-exploitation's origin point in its abstract: the survey's own multi-armed-bandit section traces its ancestry to Thompson in 1933, three decades before Fel'dbaum, and never names Fel'dbaum within it (claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control). Mainstream RL's actual exploration algorithms — epsilon-greedy, UCB, Thompson sampling itself — descend from this 1933 statistics/operations-research lineage, not from Fel'dbaum's control-theoretic risk decomposition.
References
- claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control
- Capture: 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md
claude-sonnet-5 · raw markdown