Mizutani, Dreyfus & Nishio (2000) formally derive MLP backpropagation as a special case of the Kelley-Bryson optimal-control gradient formula
The formal identity, from a paper co-authored by Stuart Dreyfus himself: MLP learning is cast as a discrete-time optimal-control problem — value function (cost-to-go), recurrence, boundary condition — and differentiating the value function yields exactly backprop's delta recursion. The paper supplies a one-to-one terminology-correspondence table between optimal-control and neural-network concepts, and situates Dreyfus's own 1962 dynamic-programming gradient derivation as producing a recursion "almost identical" to the generalized delta rule. The paper's own words for the relationship (2026-09-11 audit, direct read of the IEOR mirror): backpropagation "can be viewed as a simplified version of the Kelley-Bryson gradient formula in the classical discrete-time optimal control theory" (abstract), and the generalized delta rule "is equivalent to, what we call, the Kelley-Bryson formula" (§1). "Special case" in this note's title is the vault's gloss on "simplified version" / "equivalent", not the authors' phrase.
Contradiction candidate, flagged not resolved: this framing (backprop IS a special case of Kelley-Bryson) is stronger than Schmidhuber's (Kelley- Bryson as precursor that "lacked the efficiency" of true backprop — see claim-kelley-bryson-optimal-control-precursor). Both rest on Tier 1–2 sources; one is the historical actor's own retrospective mathematics, the other a historian's efficiency-focused reading. The disagreement is about what counts as "the same algorithm" — identity of the gradient formula vs. identity of the computational procedure — which is precisely the seam the whole priority literature turns on (claim-reverse-mode-multiple-independent-discovery).
Also carried: the paper's own novel bit (hidden-node teaching on an industrial color problem) — context only. Retrieved via the shared extraction pipeline on its second-ever bee outing. See moc-backpropagation-origins.
Source
“On derivation of MLP backpropagation from the Kelley-Bryson optimal-control gradient formula and its application”