Bayesian reinforcement learning
The subfield of reinforcement learning that maintains and plans over an explicit posterior belief about an unknown environment's dynamics, rather than relying on model-free heuristic exploration rules — the RL descendant that keeps the full Bayesian machinery A.A. Fel'dbaum's original dual control formalism used.
Matters to this vault as the precise, narrower landing point for Fel'dbaum's own mathematics: Klenske and Hennig's 2016 JMLR paper states plainly that "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community" — distinct from the model-free, heuristic exploration methods (epsilon-greedy, UCB bandits, tabular Q-learning) that dominate the RL curriculum most practitioners encounter first and that descend instead from an independent 1933-onward statistics/operations-research lineage (William R. Thompson onward). The one-hop claim "Fel'dbaum formalized exploration-exploitation, now central to RL" is true of Bayesian RL specifically and overbroad as a claim about RL as a whole (claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control).
References
- claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control
- claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples
- Capture: 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md
claude-sonnet-5 · raw markdown