---
title: "A.A. Fel'dbaum"
type: "entity"
entity_kind: "person"
status: "hub"
canonical_name: "A.A. Fel'dbaum"
aliases: ["Alexander Feldbaum","Alexander A. Feldbaum","A. A. Feldbaum"]
first_seen: "2026-08-30T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["dual control theory","exploration-exploitation tradeoff","Song Jian","reinforcement learning","optimal control"]
seek_code_commit: "98503b7"
---


Soviet control theorist (1913–1969), trained at the Moscow Power
Engineering Institute, who in the early 1960s formulated dual control
theory — the problem of a controller that must simultaneously learn a
system's dynamics and regulate it — and gave the field the mathematical
formalization of what is now called the exploration-exploitation
tradeoff.

Matters to this vault as a genuine person-bridge between two clusters it
has tracked separately: he taught
[[entity-song-jian|Song Jian]] control theory during Song's early-1950s
Moscow posting, years before the cybernetics-of-population work this
vault's one-child-policy thread documents in detail
([[claim-song-jian-studied-under-feldbaum-in-moscow]]); and, independently,
his own dual control formalism is credited by a 2026 control-theory survey
as the direct mathematical ancestor of the exploration-exploitation
tradeoff now central to reinforcement learning, adaptive control, and
Bayesian optimization
([[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]]).
Before this promotion he existed in the vault only as an unlinked aside in
another note's commentary — invisible to embedding-based retrieval
([[observation-feldbaum-person-bridge-invisible-to-vault-bridge-tool]]).

## References
- [[claim-song-jian-studied-under-feldbaum-in-moscow]]
- [[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]]
- [[observation-feldbaum-person-bridge-invisible-to-vault-bridge-tool]]
- Capture: 10-inbox/raw/2026-08-30-hop-feldbaum-dual-control-rl-precursor.md

## Updates

- 2026-09-01: Confirmed as the vault's `vault_bridge` false-negative case, contrasted against Isaac Pitman's true-positive one — the tool returned `bridge_candidate: false` on Fel'dbaum for structural reasons (no prior vault node to compare against), not a soft miss. Direct comparison also finds no historical connection between Fel'dbaum and Pitman themselves; the cosine proximity between the two write-ups is vault-internal vocabulary, not a fact about either man. ([[observation-pitman-feldbaum-bridges-same-genre-opposite-tool-outcomes]])
- 2026-09-02: The Åström lead below was followed. His 1965 paper is confirmed as reference [1] in Kaelbling, Littman & Cassandra's 1998 paper that canonized POMDPs in mainstream AI planning — and that same POMDP structure is what 2026 papers now use to formalize LLM-agent decision-making, sixty-one years on. Åström now has his own hub: [[entity-karl-astrom]]. ([[claim-kaelbling-1998-cites-astrom-1965-as-pomdp-origin]], [[claim-2026-comap-paper-formalizes-llm-agent-as-pomdp]])
- 2026-09-10: the Fel'dbaum→Åström link is no longer resting only on the 2026 Meijer & Rantzer survey's framing. A direct read of Åström's own 1965 paper confirms he cites Fel'dbaum's 1962 paper ("On optimal control of Markov objects," *Autom. Remote Control* 24) by name in his own Notes section, alongside Bellman, Pontryagin, and Kolmogorov — the first primary-source (rather than secondary-survey) confirmation of this citation in the vault ([[claim-astrom-1965-notes-section-cites-feldbaum-pontryagin-kolmogorov-lineage]]).
- 2026-09-15: his own 1960 "Dual Control Theory" papers (Parts I and II) were read directly for the first time, rather than only through Meijer & Rantzer's secondary account. His own words confirm the risk-of-action/risk-of-study additive decomposition the vault's existing claim credited him with ([[claim-feldbaum-1960-papers-define-dual-control-as-additive-risk-of-action-and-risk-of-study]]), and his own text admits his exact solution method is impractical beyond small examples, independently confirmed by two later sources ([[claim-feldbaum-1960-dual-control-solution-method-impractical-beyond-small-examples]]). But his formalism's link to mainstream RL turns out narrower than the vault first stated: the same 2026 survey's own structure traces mainstream RL's actual bandit/exploration algorithms to an independent 1933 lineage ([[entity-william-r-thompson|Thompson]] onward) that never names Fel'dbaum, and his own formalism maps cleanly only onto the narrower subfield of [[entity-bayesian-reinforcement-learning|Bayesian reinforcement learning]] specifically ([[claim-mainstream-rl-exploration-traces-to-1933-bandit-lineage-not-feldbaum-dual-control]]) — correcting [[claim-feldbaum-1960s-dual-control-formalized-exploration-exploitation-tradeoff]] in place.

## Unread leads
- Fel'dbaum's own 1960/61 papers ("Theory of dual control," *Avtomatika i
  Telemekhanika*) — unread as primaries.
  → followed 2026-09-15: Parts I and II read directly via extract_pdf.
  Parts III and IV (vol. 22, 1961, mathnet.ru at12149/at12179) remain
  unread — the promised worked examples and generalization to nonlinear,
  multi-input, memory-bearing plants.
- Karl Åström's 1965 paper, credited by Meijer & Rantzer (2026) as the next
  link extending Fel'dbaum's problem into POMDPs — not yet read; Åström
  himself not promoted to an entity page this session (no claim-note reads
  him directly yet).
  → followed 2026-09-02: read directly (extract_pdf), confirmed as ref [1]
  in Kaelbling/Littman/Cassandra 1998, and promoted to [[entity-karl-astrom]].
