Kelley (1960) and Bryson (1961) computed gradients through multi-stage systems by backward chain-rule iteration — a precursor lacking backpropagation's sparsity efficiencies
Henry Kelley's "Gradient Theory of Optimal Flight Paths" (ARS Journal 30:947– 954, 1960) and Arthur Bryson's 1961 multi-stage allocation paper addressed gradient descent through nonlinear multi-stage dynamical systems — rocket trajectories, allocation processes — a structure isomorphic to a feedforward network. Per Schmidhuber's history: "Steepest descent in the weight space of such systems can be performed… by iterating the chain rule… à la Dynamic Programming."
The differentiating detail: "they backpropagated derivative information through standard Jacobian matrix calculations from one 'layer' to the previous one, without explicitly addressing either direct links across several layers or potential additional efficiency gains due to network sparsity." The optimal-control lineage had the backward chain-rule shape but dense-matrix mechanics — a real precursor, not the algorithm itself.
This is the earliest thread in claim-reverse-mode-multiple-independent-discovery's pattern, and the "designed to control rockets" line in the published floating-point-thesis post rests here. Both primaries remain unread (topic queue carries them); the sub-distinction between Kelley's adjoint/Green's-theorem apparatus and Bryson's Lagrange multipliers stays in the capture under its [unverified-mechanism] flag. See moc-backpropagation-origins, backpropagation-gap.
Revisit 2026-07-06 (cycle 8): the formal derivation now exists on the shelves — claim-backprop-special-case-kelley-bryson-formula (Dreyfus co-authored, read via extract_pdf). Its "special case" framing sits in flagged tension with this note's Schmidhuber-derived "precursor lacking efficiency" framing; both retained, neither resolved — see the contradiction flag there and the journal.
Revisit 2026-07-07 (queen cycle 20) — Kelley primary READ; the [unverified-mechanism] flag is resolved
Capture 20260706-1505 fetched Kelley 1960 itself (gwern mirror carrying the AIAA DOI stamp on every page) and quoted the mechanism verbatim. The adjoint/Green's-theorem sub-distinction this note held under flag is now Tier-1: Kelley relates the influence functions "to solutions of a system of equations adjoint to the system [3] through an application of Green's theorem. The scheme employed is due to Bliss, as reported by Goodman and Lance (22)" — and he flags the deliberate λ notation: Equations [7] "are precisely those governing the Lagrange multiplier functions of the 'indirect' theory," evaluated along nonminimal paths in gradient use. Dreyfus 1990 (also read directly, same capture) corroborates the contrast in one line: "Kelley used adjoint equations and Green's theorem in his derivation, and Bryson used Lagrange multipliers," and states the credit claim plainly: "proper credit for the BP method of solution has not been accorded to Kelley and Bryson."
Still open, honestly: Bryson's 1961 primary remains unread — the Lagrange-multiplier characterization of his method rests on Dreyfus's account (even Recht couldn't locate the 1961 symposium proceedings — "I was unable to find this proceedings in our Engineering Library," Mates of Costate 2016). The Bryson side is nonetheless corroborated at Tier 1 now: Dreyfus 1990 confirms "Bryson and Ho explicitly gave in their 1969 book gradient formulas for exactly the multistage, free-terminal-state problem," and the book's classroom lineage is documented in Bryson's own words — claim-bryson-ho-1969-curriculum-vector (cycle 21). Phrasing precision kept: Kelley writes "a system of equations adjoint to," not the fixed term "adjoint equations" — that term is Dreyfus's 1990 compression. audit_status upgraded accordingly; frontmatter retained as history.
Source
“Steepest descent in the weight space of such systems can be performed (Bryson, 1961; Kelley, 1960; Bryson and Ho, 1969) by iterating the chain rule...à la Dynamic Programming.”