---
title: "Bayesian reinforcement learning"
type: "entity"
entity_kind: "concept"
status: "hub"
canonical_name: "Bayesian reinforcement learning"
aliases: ["Bayesian RL"]
first_seen: "2026-09-15T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["A.A. Fel'dbaum / dual control theory","reinforcement learning","POMDP","exploration-exploitation tradeoff","dynamic programming"]
seek_code_commit: "546fa57"
---


The subfield of reinforcement learning that maintains and plans over an explicit posterior belief about an unknown environment's dynamics, rather than relying on model-free heuristic exploration rules — the RL descendant that keeps the full Bayesian machinery [[entity-aa-feldbaum|A.A. Fel'dbaum]]'s original dual control formalism used.

Matters to this vault as the precise, narrower landing point for Fel'dbaum's own mathematics: Klenske and Hennig's 2016 *JMLR* paper states plainly that "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community" — distinct from the model-free, heuristic exploration methods (epsilon-greedy, UCB bandits, tabular Q-learning) that dominate the RL curriculum most practitioners encounter first and that descend instead from an independent 1933-onward statistics/operations-research lineage ([[entity-william-r-thompson|William R. Thompson]] onward). The one-hop claim "Fel'dbaum formalized exploration-exploitation, now central to RL" is true of Bayesian RL specifically and overbroad as a claim about RL as a whole ([[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]).

## References
- [[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]
- [[claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples]]
- Capture: 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md
