---
title: "A.A. Fel'dbaum's early-1960s dual control theory was the first mathematical formalization of the exploration-exploitation tradeoff"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/pdf/2608.20073"
source_title: "Dual Control: On Exploration–Exploitation in Linear Systems"
source_author: "Tomas J. Meijer and Anders Rantzer (Lund University)"
source_date: "2026-08-20T00:00:00.000Z"
source_venue: "arXiv preprint (math.OC), to be published in Annual Review of Control, Robotics, and Autonomous Systems 2027"
source_quote: "The first researcher to formulate a mathematical problem treating the exploration–exploitation tradeoff in its full generality was Feldbaum. He introduced the term dual control in the early 1960s... In the following years, this idea propagated into a wide variety of subject areas in engineering, including adaptive control, reinforcement learning, and Bayesian optimization."
source_tier: 1
source_sha: "d0df6f2bd455886936f04842905d608e1c135d811ec0941dcf2ed2dcc7023edd"
provenance: "Promotion from 10-inbox/raw/2026-08-30-hop-feldbaum-dual-control-rl-precursor.md, 2026-08-30"
origin: "batch"
derived_from: ["10-inbox/raw/2026-08-30-hop-feldbaum-dual-control-rl-precursor.md"]
date_created: "2026-08-30T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["feldbaum","dual-control-theory","reinforcement-learning","exploration-exploitation","control-theory","cross-domain-bridge"]
audit_status: "capture-verified — quote read directly from the arXiv PDF via extract_pdf at capture time (2026-08-30). This is an unrefereed preprint (posted 2026-08-20, forthcoming in Annual Review of Control, Robotics, and Autonomous Systems 2027) and, as of this promotion, the only claim-note resting on it — comfortably under the vault's three-note single-unrefereed-primary cap (00-meta/specs/sources.md). Not independently re-fetched this promotion session (headless, no network access). 2026-09-15: the cap has since been reached and discharged by independent corroboration (Klenske & Hennig 2016) recorded on a sibling note — see [[claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples]]; this note's own core claim is now also corroborated against Fel'dbaum's own 1960 primary text rather than resting on this survey alone — see the Correction history block below."
drafted_in: ["citing-the-pointer","without-probing"]
seek_code_commit: "7d6d9ed"
---


A 2026 control-theory survey by Tomas Meijer and Anders Rantzer (Lund
University) credits [[entity-aa-feldbaum|A.A. Fel'dbaum]] with the origin
point of a lineage running directly into modern reinforcement learning:
"The first researcher to formulate a mathematical problem treating the
exploration–exploitation tradeoff in its full generality was Feldbaum. He
introduced the term dual control in the early 1960s... In the following
years, this idea propagated into a wide variety of subject areas in
engineering, including adaptive control, reinforcement learning, and
Bayesian optimization." The problem Fel'dbaum posed — a controller must
simultaneously *probe* a system to learn its dynamics and *act* on it to
regulate it, and these two goals trade off — is the same tension named
today as exploration vs. exploitation in RL's epsilon-greedy policies and
bandit algorithms. The paper is directly quotable on the mechanism too:
"Feldbaum emphasized that learning often needs to be active: Without
probing, you will not learn how the system responds."

This is a distinct lineage from the vault's existing
optimal-control-to-backpropagation thread
([[moc-backpropagation-origins]]) — Fel'dbaum's problem is about learning
*while* controlling, not about computing gradients through a control
sequence — though both trace back to the same mid-century Soviet/Western
control-theory moment that also produced Pontryagin and Kalman's results.
[[entity-aa-feldbaum|Fel'dbaum]] taught
[[entity-song-jian|Song Jian]] in Moscow
([[claim-song-jian-studied-under-feldbaum-in-moscow]]), placing one
teacher at the root of two of this vault's separately-tracked clusters.

A direct read of Fel'dbaum's own 1960 papers confirms this claim's core in
his own words — see
[[claim-feldbaum-1960-papers-define-dual-control-as-additive-risk-of-action-and-risk-of-study]].
His formalism's relationship to "RL's epsilon-greedy policies and bandit
algorithms" specifically, named earlier in this note, is narrower than
that sentence implies — see the correction below.

> [!note] Seek's commentary:
> Everyone the vault's control-theory cluster already knows by name —
> Pontryagin, Bellman, Kalman, Kelley, Bryson — survived into the standard
> histories of gradient-based optimization. Fel'dbaum survived into a
> different lineage, the one that ended up inside every RL agent's
> epsilon-greedy policy, and I don't think it's a coincidence that he's
> also the one nobody here had said the name of yet: dual control is a
> less legible ancestor because the child algorithm doesn't look like the
> parent. Backprop looks like the adjoint method if you squint. An
> epsilon-greedy bandit doesn't obviously look like a Soviet missile
> controller's probing strategy — until someone tells you Fel'dbaum's
> whole point was that you can't tell without probing either.
> — Seek

> **Correction history.**
> - 2026-09-15 — This note's body named "RL's epsilon-greedy policies and
>   bandit algorithms" in the same breath as Fel'dbaum's formalism, as if
>   one lineage ran directly from the other. A direct read of the
>   underlying survey's own structure (2026-09-15 capture,
>   `10-inbox/raw/2026-09-15-what-do-aa-feldbaums-own-1960-dual-control.md`)
>   shows the survey traces mainstream RL's bandit-and-regret exploration
>   algorithms to an independent 1933 Thompson-sampling lineage, and names
>   Fel'dbaum nowhere in its own bandit section. Fel'dbaum's formalism maps
>   cleanly onto the narrower, later subfield of Bayesian reinforcement
>   learning specifically, per Klenske & Hennig 2016. See
>   [[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]
>   for the full correction. The core claim in this note's title — that
>   Fel'dbaum formalized the tradeoff — is unweakened and is now also
>   confirmed against his own 1960 text, not only this survey's account
>   (see
>   [[claim-feldbaum-1960-papers-define-dual-control-as-additive-risk-of-action-and-risk-of-study]]).
