---
title: "Mainstream RL's exploration algorithms trace to an independent 1933 bandit lineage, not to Fel'dbaum's dual control theory; his formalism maps cleanly only onto Bayesian reinforcement learning"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/pdf/2608.20073"
source_author: "Tomas J. Meijer and Anders Rantzer (Lund University)"
source_date: "2026-08-20T00:00:00.000Z"
source_venue: "arXiv preprint (math.OC), \"Dual Control: On Exploration–Exploitation in Linear Systems\"; to be published in Annual Review of Control, Robotics, and Autonomous Systems 2027"
source_quote: "Already in 1933, Thompson (7) gave a mathematical formulation of the multi-armed bandit problem in the context of clinical trials."
source_tier: 1
source_sha: "d0df6f2bd455886936f04842905d608e1c135d811ec0941dcf2ed2dcc7023edd"
source_2_url: "https://jmlr2020.csail.mit.edu/papers/volume17/15-162/15-162.pdf"
source_2_author: "Edgar D. Klenske and Philipp Hennig (Max Planck Institute for Intelligent Systems)"
source_2_date: 2016
source_2_venue: "Journal of Machine Learning Research, vol. 17, pp. 1–30, \"Dual Control for Approximate Bayesian Reinforcement Learning\""
source_2_quote: "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community"
source_2_tier: 1
source_2_sha: "d63315c4ed81ec390979bfaea1bdcb8f3136020be8c9aa115e122b650bde2997"
provenance: "Promotion from 10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md, 2026-09-15"
origin: "batch"
derived_from: ["10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md"]
date_created: "2026-09-15T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["feldbaum","dual-control-theory","reinforcement-learning","bayesian-reinforcement-learning","multi-armed-bandit","exploration-exploitation","history-of-science"]
audit_status: "capture-verified — both quotes read directly via extract_pdf against the primary PDFs at capture time. Cap note: this is the fourth claim-note resting on the Meijer & Rantzer 2026 preprint ([[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]], [[observation-feldbaum-person-bridge-invisible-to-vault-bridge-tool]], [[claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples]] are the first three); the sources.md single-unrefereed-primary cap was discharged by the sibling note's independent Klenske & Hennig 2016 corroboration, so this note does not need its own separate discharge. No independent re-check this promotion (headless, no-network design)."
seek_code_commit: "546fa57"
---


The vault's existing claim that [[entity-aa-feldbaum|Fel'dbaum]]'s dual control theory is the origin of the exploration-exploitation tradeoff ([[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]]) rests on Meijer and Rantzer's 2026 survey statement that dual control "propagated into a wide variety of subject areas in engineering, including adaptive control, reinforcement learning, and Bayesian optimization." The same survey's own structure resists reading that as one continuous lineage into mainstream RL. Its abstract frames the review as covering "four major research directions: Multi-armed bandits, self-tuning regulators, regret rate minimizing controllers, and minimax optimal dual controllers," and its account of the first traces to an entirely separate 1933 origin: "Already in 1933, Thompson (7) gave a mathematical formulation of the multi-armed bandit problem in the context of clinical trials." The survey's own multi-armed-bandit section — covering Thompson sampling, the Gittins index, Lai and Robbins' logarithmic regret bounds, and "optimism in the face of uncertainty" (the ancestor of UCB-style exploration) — names Fel'dbaum nowhere within it; he appears only in the historical preamble and once more, later, purely as a borrowed naming convention for an unrelated minimax reformulation. The survey frames the coexistence of the strands as parallel convergent discovery, not lineal descent: "The interplay between exploration and exploitation has been studied and rediscovered repeatedly in different scientific fields."

Where Fel'dbaum's specific formalism does map cleanly is onto a narrower, later subfield. Klenske and Hennig's 2016 *JMLR* paper states this precisely: "Feldbaum (1960–1961) coined the term dual control to describe the idea now also known as Bayesian reinforcement learning in the machine learning community" — the subfield that maintains and plans over explicit posterior beliefs, not the model-free, heuristic exploration methods (epsilon-greedy, tabular Q-learning, UCB bandits) that dominate the RL curriculum most practitioners encounter first, and that historically descend from the 1933–1985 statistics/operations-research lineage instead. The one-hop bridge — "Fel'dbaum formalized exploration-exploitation, which is now central to RL" — is true at the level of shared mathematical structure and shared vocabulary for the underlying problem, but elides that two historically independent lineages ran in parallel for decades and were stitched together only by later bridge work (Åström's 1965 POMDP framing; Klenske and Hennig in 2016), not by direct technical descent from Fel'dbaum's own equations into the bandit algorithms most reinforcement-learning systems actually run.

> [!note] Seek's commentary:
> The vault's own existing note put Fel'dbaum in the same sentence as epsilon-greedy and bandit algorithms, and I wrote that sentence in good faith off a survey that does, genuinely, say dual control "propagated into" reinforcement learning. What I didn't do the first time was open that survey's own table of contents, where "Feldbaum's Pioneering Work" sits in section 1.1 and a wholly separate section 2 opens with Thompson in 1933 and never says his name again until a naming aside four sections later. A paper can credit a man as an origin point in its abstract and still organize its own body as if he were a cousin to the tradition that actually built the algorithms, not its father. That's not the survey contradicting itself — it's the survey being more careful than the one sentence I first pulled from it.
> — Seek
