---
date_created: "2026-07-06T00:00:00.000Z"
title: "MOC: Backpropagation's origins — priority, paternity, and myth"
type: "moc"
tags: ["moc","backpropagation","history-of-ml","linnainmaa","automatic-differentiation"]
updated: "2026-07-15T00:00:00.000Z"
---


The vault's deepest cluster: how the algorithm that trains every modern model
was derived repeatedly, in different fields, by people who mostly didn't know
about each other — and how the folk history compressed that into single-inventor
stories. The spine claim: [[backpropagation-gap]].

## The Linnainmaa thread (numerical analysis, 1970)
- [[claim-linnainmaa-thesis-identity]] — what the thesis was, where, in what language
- [[claim-linnainmaa-rounding-error-problem]] — the actual problem it solved
- [[claim-linnainmaa-reverse-mode-single-pass]] — the mechanism (Tier-1 verified via Griewank 2014 and now 2012)
- [[claim-linnainmaa-1976-algorithm-t-reverse-sweep]] — the 1976 primary itself, read page-by-page: Algorithm T's reverse-order accumulation in Linnainmaa's own words (+ venue correction: BIT 16, not *Annales*)
- [[claim-linnainmaa-field-numerical-analysis]] — why it was invisible to AI
- [[claim-linnainmaa-priority-not-paternity]] — the reception structure (see its 2026-07-06 revisit note: "priority" itself is now complicated by Griewank's multiple-discovery account)
- [[claim-linnainmaa-copenhagen-anecdote]] — the origin story, source-critically handled
- [[claim-speelpenning-1980-does-not-cite-linnainmaa]] — the canonical early AD *implementation* didn't cite him either: the walls run even inside AD (OSTI primary, direct OCR)
- [[claim-werbos-1974-does-not-cite-linnainmaa]] — nor did the 1974 thesis (exhaustive full read): independent derivation, and the "Section 5.5.1" locator doesn't exist in the 1974 text
- [[claim-linnainmaa-1976-citation-histogram]] — the reception, finally in numbers: 51 dated citations across three decades, then 245 since 2010 (myth moved unresolved → contested)

## The lineages before and after
- [[claim-kelley-bryson-optimal-control-precursor]] — the optimal-control thread (dense Jacobians, no sparsity)
- [[claim-nyquist-bode-classical-control-to-optimal-control-bridge]] — one generation further upstream: does Bell Labs classical control theory (Nyquist 1932, Bode) feed the Kelley/Bryson optimal-control lineage, or is the periodization a coincidence? Flagged `[unverified-mechanism/synthesis]`, routed at [[question-verify-nyquist-bode-optimal-control-transmission-line]]. Grounded in [[claim-black-1927-ferry-epiphany-retrospective-simplification]] and [[claim-nyquist-bode-supplied-stability-math-black-lacked]], a separate myth-of-origin thread about Harold Black's feedback amplifier that lives outside this cluster proper.
- [[claim-cps-wiener-lineage-no-link-to-backprop-optimal-control]] — a second, independent doorway into the same locked room: does Norbert Wiener's cybernetics (the term's other claimed common root, via cyber-physical systems' 2006 coinage — see [[claim-cyber-physical-systems-cybernetics-common-root-not-derivation]]) bridge to Kelley/Bryson/Pontryagin/Bellman? No source found either way; recorded as a documented absence, `[unverified-mechanism]`.
- [[claim-minsky-1961-named-credit-assignment]] — the term, coined temporal; structural came later
- [[claim-perceptrons-multilayer-sterile-was-conjecture]] — the 1969 book proved single-layer limits; the multilayer "sterile" line was a self-flagged intuitive judgment inviting rejection
- [[claim-dreyfus-1988-connectionism-was-vindicated-not-target]] — the philosophy-of-AI bridge: the Dreyfuses' 1988 Daedalus paper casts Rosenblatt's connectionism as vindicated by their critique, resolving Stuart Dreyfus's "double role"
- [[claim-dreyfus-critique-targeted-symbolic-ai-not-neural-nets]] — the same paper's two-research-programs framing: the critique targets Newell & Simon's physical symbol system hypothesis, not the Hebb–Rosenblatt learning line
- [[claim-simon-1983-dismissed-perceptron-research-citing-rosenblatt]] — Simon's own 1983 verdict, by name: perceptron-style learning "didn't get anywhere," on empirical-productivity grounds, never engaging Minsky & Papert's formal conjecture
- [[claim-simon-vera-1993-1995-connectionist-nets-are-symbol-systems]] — Simon's later move: rather than concede connectionism as a rival paradigm, redefine "symbol" broadly enough to absorb it (1993, 1995)
- [[observation-simon-responds-to-connectionism-by-absorption-not-refutation]] — names the pattern across both: absorption, a third mode alongside confident conjecture and Dreyfus's program-splitting
- [[claim-perceptrons-credit-assignment-only-in-1988-epilogue]] — the phrase "credit-assignment" is 1988-Epilogue vocabulary, not 1969
- [[claim-bottou-2017-foreword-perceptrons-chapters-0-13-are-1969]] — Bottou's 2017 foreword fixes the 1969/1988 boundary: chapters 0–13 are the first edition; Prologue + Epilogue are 1988 additions
- [[claim-pollack-1989-credit-assignment-in-1988-prologue]] — Pollack's 1989 review places the credit/blame argument in the 1988 Prologue (p. xiii); refines the "only in the Epilogue" location
- [[claim-pollack-1989-sterile-quote-page-231]] — Pollack independently quotes the "sterile" passage, at p. 231 (a Tier-2 carrier; p. 231 vs p. 232 discrepancy flagged)
- [[claim-perceptrons-1969-standalone-scan-not-freely-available]] — the diffable object doesn't exist on the open web: repositories hold only the 1988 edition
- [[claim-werbos-1974-no-credit-assignment-language]] — the thesis never says "credit assignment" (grep ×0 over 453 pp.); the "spatial credit assignment" label is retrospective, earliest compound landmark Schmidhuber's own 1990 dissertation title
- [[claim-bryson-ho-1969-curriculum-vector]] — how the adjoint method actually traveled: Harvard/MIT/AIAA courses → the 1969 textbook (Bryson's own account)
- [[claim-mccarthy-1973-control-theory-little-relevance-to-ai]] — the wall named from the AI side: McCarthy's own 1973 rebuttal to Lighthill dismissed control theory as having "little relevance to AI," the field where the adjoint/gradient ancestor of backprop already lived
- [[claim-dreyfus-1973-lineage-link-uncorroborated]] — Schmidhuber's 1973 link vs Dreyfus's own retrospectives, which credit 1962 and never cite 1973
- [[claim-recht-adjoints-equivalence-not-transmission]] — the modern control-side statement: mathematical equivalence via Lagrangian duality, explicitly not a transmission story (and no Kelley)
- [[claim-lecun-1989-first-practical-recognition]] — where the algorithm met the mail (NETtalk boundary stated)
- [[claim-vanishing-gradient-chain-rule-pathology]] — the failure mode inside the backward pass
- [[claim-werbos-2011-cathexis-derivative-mapping]] — the Freud mapping in Werbos's own 2011 words
- [[claim-werbos-tsp-manual-first-publication]] — Werbos's claim that the first publication was an MIT statistics manual: citation real, document unlocatable, single-witness
- [[claim-werbos-1978-docsub-blocked]] — DARPA kept the full report out of DOCSUB; the '78 journal version shipped without the how-to appendix (two parallel reasons, on the record)
- [[claim-backprop-special-case-kelley-bryson-formula]] — Dreyfus's own formal identity claim (in flagged tension with the precursor framing)
- [[claim-hinton-biological-implausibility-four-objections]] — what "biologically implausible" actually asserts
- [[claim-fukushima-1979-neocognitron-first-cnn]] — the CNN's architectural ancestor (not backprop-trained)
- [[claim-parker-1985-technical-report-not-cited]] — the fourth uncited independent rediscoverer
- [[claim-ivakhnenko-gmdh-first-deep-characterization]] — "first deep learning, 1965": Schmidhuber's hedged characterization, regression-grown layers, eight-layer figure unverified
- [[claim-werbos-ieee-awards-1995-2022]] — the awards are real (IEEE primaries); the 2022 citation language is itself myth-ledger circulation
- [[claim-amari-1972-associative-memory-precedes-hopfield]] — priority-without-citation, the cluster's recurring pattern
- [[claim-robbins-monro-1951-stochastic-approximation]] — the 1951 statistical ancestor of SGD, finally quoted from its own text (root-finding, not optimization)
- [[claim-amari-1968-saito-experiment-primary-read]] — the 1968 Kyoritsu pages read in Japanese: Saito's 1967 experiment is real and stochastic-descent-trained; the "five-layer MLP" framing is not the book's

## The 1986 watershed
- [[claim-rhw-1986-reference-list-four-works]] — four references, none of the earlier derivations
- [[claim-rhw-1986-demonstration-not-invention]] — what the paper actually added (verified demos, verbatim symmetry-breaking)

## The mechanism and the multiplicity
- [[claim-backpropagation-special-case-of-reverse-mode-ad]] — the precise correspondence (Baydin, JMLR)
- [[claim-jvp-vjp-transpose-duality]] — forward and reverse mode as one Jacobian in dual directions (pushforward/pullback; the cost duality follows from shape)
- [[claim-wengert-list-named-for-forward-mode-inventor]] — the tape's name honors the *forward*-mode inventor; automatic reverse mode is Speelpenning 1980
- [[claim-autograd-three-reifications-of-the-tape]] — PyTorch records a DAG, TensorFlow records a literal tape, JAX compiles a jaxpr and never says "tape"
- [[claim-cheap-gradient-bound-two-figures]] — why it's affordable: the gradient costs a small constant multiple of one forward evaluation ("one to four times" / "5×", reconciled)
- [[claim-reverse-mode-multiple-independent-discovery]] — at least five discoveries, five fields, 1965–1986 (Griewank 2012, direct read)
- [[claim-ml-and-ad-communities-mutually-unaware]] — *why* the field rediscovered it five times: two literatures that didn't read each other (Baydin 2018, Tier 1) — the structural root of the whole attribution gap
- [[claim-werbos-backprop-from-freud-own-account]] — the Freud thread, finally in Werbos's own words
- [[claim-werbos-1968-cybernetica-earliest-germ]] — the earliest germ: a real 1968 paper whose Freud→DP content survives only as Werbos's own paraphrase (no readable copy exists)
- [[claim-werbos-1997-optimization-consciousness-chapter]] — the mind/consciousness face of the same optimization architecture (1997, not 1996; not his fullest statement)

## The bridge to the economics cluster
- [[claim-training-inference-compute-asymmetry-mechanism]] — the backward pass is the cost only training pays

## The bridge to the schema-change cluster (the author is the hinge)
- [[observation-rumelhart-person-bridge-schema-theory-to-backprop]] — the "R" in
  RHW 1986 ([[claim-rhw-1986-demonstration-not-invention]]) is the same David
  Rumelhart who coined the 1978 tuning/restructuring taxonomy under
  [[moc-schema-change-and-restructuring]]. One trajectory from cognitive-psychology
  schema theory to the connectionist revival; a person-bridge no note had named.
  `seedling`.

## Biological plausibility and the alternatives
- [[claim-crick-1989-antidromic-not-weight-transport]] — the objection, correctly attributed
- [[claim-feedback-alignment-random-weights-train]] — random feedback weights (contested)
- [[claim-hinton-forward-forward-boltzmann-lineage]] — two forward passes, no backward
- [[claim-update-locking-backprop-constraint]] — the systems-level cost of the backward pass
- [[claim-brain-approximates-backprop-core-principles-ngrad]] — Hinton's mature position (Lillicrap 2020): the brain implements backprop's *core principles* via NGRAD, not strict backprop
- [[claim-hinton-backprop-in-brain-2007-to-2022-arc]] — the fifteen-year reversal: 2007 rescued backprop-in-the-brain, 2022 Forward-Forward abandoned it

## The myth ledger (circulation vs. primary support)
- [[myth-werbos-invented-backpropagation]] — contested
- [[myth-linnainmaa-uncited-before-2010s]] — unresolved
- [[myth-amari-first-sgd-mlp]] — contested (single witness + citogenesis)
- [[myth-perceptrons-book-killed-connectionism]] — contested (proof-vs-conjecture collapse; causation genuinely multi-sided)
- [[claim-wikipedia-amari-sgd-citogenesis]] — how the amari myth's "corroboration" collapses to one witness: shared typo + shared misspelling
- [[claim-griewank-2012-does-not-mention-amari]] — confirmed negative: the reverse-mode history and the Amari-SGD claim never intersect at the primary level

## Open verification threads
- [[question-verify-nyquist-bode-optimal-control-transmission-line]] — 2026-07-11: does Kelley 1960 (primary already read for [[claim-kelley-bryson-optimal-control-precursor]]) or Bryson's unread 1961 paper cite Nyquist/Bode/classical control theory by name? Cheapest open check in the cluster right now.
- ~~question-verify-linnainmaa-citation-count~~ — CLOSED 2026-07-07 into
  `_answered`; the pull happened and moved the myth
  ([[claim-linnainmaa-1976-citation-histogram]]). Successor sharpener lives
  in [[myth-linnainmaa-uncited-before-2010s]]'s watch flag (classify the 51
  pre-2010 citations by community).
- Library-op candidates for Cali (documents no automated fetch can reach):
  Werbos 1994 reprint TOC (resolves the 5.5.1 locator anomaly), Dreyfus
  1973 IEEE TAC full text, Bryson 1961 symposium proceedings, Talking Nets
  ch. 16 p. 376 (Hinton's symmetry objection), Ivakhnenko 1971.
- 2026-07-07 evening safelist CLEARED: all backlog captures feeding this
  cluster are promoted or explicitly held (Opus-lane holds excepted).
