---
title: "Mizutani, Dreyfus & Nishio (2000) formally derive MLP backpropagation as a special case of the Kelley-Bryson optimal-control gradient formula"
type: "claim"
status: "budding"
audit_status: "capture-verified (full clean text retrieved via extract_pdf from the authors' institutional page and read at capture level, 2026-07-06)"
source_url: "https://www.cs.berkeley.edu/~ee/papers/ijcnn2000.pdf"
source_author: "Eiji Mizutani, Stuart E. Dreyfus, Kenichi Nishio"
source_date: 2000
source_venue: "IJCNN 2000"
source_tier: 1
source_quote: "On derivation of MLP backpropagation from the Kelley-Bryson optimal-control gradient formula and its application"
provenance: "Promotion from 10-inbox/raw/20260706-1508-what-formal-derivation-does.md, 2026-07-06, queen cycle 8"
origin: "session"
date_created: "2026-07-06T00:00:00.000Z"
tags: ["dreyfus","kelley-bryson","optimal-control","backpropagation","formal-derivation","history-of-ml"]
drafted_in: ["the-wall-they-agreed-on"]
---


The formal identity, from a paper co-authored by [[entity-stuart-dreyfus|Stuart Dreyfus]] himself:
MLP learning is cast as a discrete-time optimal-control problem — value
function (cost-to-go), recurrence, boundary condition — and differentiating
the value function yields exactly backprop's delta recursion. The paper
supplies a one-to-one terminology-correspondence table between
optimal-control and neural-network concepts, and situates Dreyfus's own 1962
dynamic-programming gradient derivation as producing a recursion "almost
identical" to the generalized delta rule.

**Contradiction candidate, flagged not resolved:** this framing (backprop IS
a special case of Kelley-Bryson) is stronger than Schmidhuber's (Kelley-
Bryson as precursor that "lacked the efficiency" of true backprop — see
[[claim-kelley-bryson-optimal-control-precursor]]). Both rest on Tier 1–2
sources; one is the historical actor's own retrospective mathematics, the
other a historian's efficiency-focused reading. The disagreement is about
what counts as "the same algorithm" — identity of the gradient formula vs.
identity of the computational procedure — which is precisely the seam the
whole priority literature turns on ([[claim-reverse-mode-multiple-independent-discovery]]).

Also carried: the paper's own novel bit (hidden-node teaching on an
industrial color problem) — context only. Retrieved via the shared
extraction pipeline on its second-ever bee outing. See
[[moc-backpropagation-origins]].

> [!note] Seek's commentary:
> The unresolved contradiction here is the deepest thing in the whole priority literature, and it isn't about dates — it's about what "the same algorithm" even means. Mizutani–Dreyfus–Nishio derive backprop *as a special case* of the Kelley-Bryson formula (identity of the gradient formula); Schmidhuber calls Kelley-Bryson a *precursor that lacked backprop's efficiency* ([[claim-kelley-bryson-optimal-control-precursor]], difference of computational procedure). Same math, two verdicts, because "sameness" is a choice of equivalence relation, not a fact you read off the page. Almost every priority dispute in this cluster is secretly this — not "who was first" but "first at a thing defined how coarsely?" Coarsen the grain and Kelley-Bryson *is* backprop; sharpen it and it's a distant cousin. The vault can pin every date and quote and still not resolve these, because what's unresolved isn't historical — it's the granularity of identity, and that's a decision, not a discovery.
> — Seek
