MOC: Backpropagation's origins — priority, paternity, and myth
The vault's deepest cluster: how the algorithm that trains every modern model was derived repeatedly, in different fields, by people who mostly didn't know about each other — and how the folk history compressed that into single-inventor stories. The spine claim: backpropagation-gap.
The Linnainmaa thread (numerical analysis, 1970)
- claim-linnainmaa-thesis-identity — what the thesis was, where, in what language
- claim-linnainmaa-rounding-error-problem — the actual problem it solved
- claim-linnainmaa-reverse-mode-single-pass — the mechanism (Tier-1 verified via Griewank 2014 and now 2012)
- claim-linnainmaa-1976-algorithm-t-reverse-sweep — the 1976 primary itself, read page-by-page: Algorithm T's reverse-order accumulation in Linnainmaa's own words (+ venue correction: BIT 16, not Annales)
- claim-linnainmaa-field-numerical-analysis — why it was invisible to AI
- claim-linnainmaa-priority-not-paternity — the reception structure (see its 2026-07-06 revisit note: "priority" itself is now complicated by Griewank's multiple-discovery account)
- claim-linnainmaa-copenhagen-anecdote — the origin story, source-critically handled
- claim-speelpenning-1980-does-not-cite-linnainmaa — the canonical early AD implementation didn't cite him either: the walls run even inside AD (OSTI primary, direct OCR)
- claim-werbos-1974-does-not-cite-linnainmaa — nor did the 1974 thesis (exhaustive full read): independent derivation, and the "Section 5.5.1" locator doesn't exist in the 1974 text
- claim-linnainmaa-1976-citation-histogram — the reception, finally in numbers: 51 dated citations across three decades, then 245 since 2010 (myth moved unresolved → contested)
The lineages before and after
- claim-isaacs-1951-rand-paper-precursors-maximum-principle — a third, earlier node: Rufus Isaacs's first RAND paper (1951) already held precursors of the maximum principle, dynamic programming, and backward analysis, five years before Pontryagin's 1956 announcement — a historian's characterization (Breitner 2005, Tier 2), primary itself unread.
- claim-bellman-named-dynamic-programming-as-political-camouflage — the term's own etymology, in Bellman's own words: "dynamic programming" was chosen as bureaucratic camouflage against a Secretary of Defense who feared the word "research," not as a description of the method. A term this cluster treats as a technical constant turns out to be a political artifact.
- claim-kelley-bryson-optimal-control-precursor — the optimal-control thread (dense Jacobians, no sparsity)
- claim-nyquist-bode-classical-control-to-optimal-control-bridge — one generation further upstream: does Bell Labs classical control theory (Nyquist 1932, Bode) feed the Kelley/Bryson optimal-control lineage, or is the periodization a coincidence? Flagged
[unverified-mechanism/synthesis], routed at question-verify-nyquist-bode-optimal-control-transmission-line. Grounded in claim-black-1927-ferry-epiphany-retrospective-simplification and claim-nyquist-bode-supplied-stability-math-black-lacked, a separate myth-of-origin thread about Harold Black's feedback amplifier that lives outside this cluster proper. - claim-cps-wiener-lineage-no-link-to-backprop-optimal-control — a second, independent doorway into the same locked room: does Norbert Wiener's cybernetics (the term's other claimed common root, via cyber-physical systems' 2006 coinage — see claim-cyber-physical-systems-cybernetics-common-root-not-derivation) bridge to Kelley/Bryson/Pontryagin/Bellman? No source found either way; recorded as a documented absence,
[unverified-mechanism]. - claim-minsky-1961-named-credit-assignment — the term, coined temporal; structural came later
- claim-perceptrons-multilayer-sterile-was-conjecture — the 1969 book proved single-layer limits; the multilayer "sterile" line was a self-flagged intuitive judgment inviting rejection
- claim-dreyfus-1988-connectionism-was-vindicated-not-target — the philosophy-of-AI bridge: the Dreyfuses' 1988 Daedalus paper casts Rosenblatt's connectionism as vindicated by their critique, resolving Stuart Dreyfus's "double role"
- claim-dreyfus-critique-targeted-symbolic-ai-not-neural-nets — the same paper's two-research-programs framing: the critique targets Newell & Simon's physical symbol system hypothesis, not the Hebb–Rosenblatt learning line
- claim-simon-1983-dismissed-perceptron-research-citing-rosenblatt — Simon's own 1983 verdict, by name: perceptron-style learning "didn't get anywhere," on empirical-productivity grounds, never engaging Minsky & Papert's formal conjecture
- claim-simon-vera-1993-1995-connectionist-nets-are-symbol-systems — Simon's later move: rather than concede connectionism as a rival paradigm, redefine "symbol" broadly enough to absorb it (1993, 1995)
- observation-simon-responds-to-connectionism-by-absorption-not-refutation — names the pattern across both: absorption, a third mode alongside confident conjecture and Dreyfus's program-splitting
- claim-perceptrons-credit-assignment-only-in-1988-epilogue — the phrase "credit-assignment" is 1988-Epilogue vocabulary, not 1969
- claim-bottou-2017-foreword-perceptrons-chapters-0-13-are-1969 — Bottou's 2017 foreword fixes the 1969/1988 boundary: chapters 0–13 are the first edition; Prologue + Epilogue are 1988 additions
- claim-pollack-1989-credit-assignment-in-1988-prologue — Pollack's 1989 review places the credit/blame argument in the 1988 Prologue (p. xiii); refines the "only in the Epilogue" location
- claim-pollack-1989-sterile-quote-page-231 — Pollack independently quotes the "sterile" passage, at p. 231 (a Tier-2 carrier; p. 231 vs p. 232 discrepancy flagged)
- claim-perceptrons-1969-standalone-scan-not-freely-available — the diffable object doesn't exist on the open web: repositories hold only the 1988 edition
- claim-werbos-1974-no-credit-assignment-language — the thesis never says "credit assignment" (grep ×0 over 453 pp.); the "spatial credit assignment" label is retrospective, earliest compound landmark Schmidhuber's own 1990 dissertation title
- claim-bryson-ho-1969-curriculum-vector — how the adjoint method actually traveled: Harvard/MIT/AIAA courses → the 1969 textbook (Bryson's own account)
- claim-mccarthy-1973-control-theory-little-relevance-to-ai — the wall named from the AI side: McCarthy's own 1973 rebuttal to Lighthill dismissed control theory as having "little relevance to AI," the field where the adjoint/gradient ancestor of backprop already lived
- claim-dreyfus-1973-lineage-link-uncorroborated — Schmidhuber's 1973 link vs Dreyfus's own retrospectives, which credit 1962 and never cite 1973
- claim-recht-adjoints-equivalence-not-transmission — the modern control-side statement: mathematical equivalence via Lagrangian duality, explicitly not a transmission story (and no Kelley)
- claim-lecun-1989-first-practical-recognition — where the algorithm met the mail (NETtalk boundary stated)
- claim-vanishing-gradient-chain-rule-pathology — the failure mode inside the backward pass
- claim-werbos-2011-cathexis-derivative-mapping — the Freud mapping in Werbos's own 2011 words
- claim-werbos-tsp-manual-first-publication — Werbos's claim that the first publication was an MIT statistics manual: citation real, document unlocatable, single-witness
- claim-werbos-1988-tsp-manual-citation-dated-1975 — the manual's date, per Werbos's own 1988 bibliography, repeated three times, never 1977
- claim-klensin-romberg-1989-sole-source-of-tsp-manual-1977-date — where the vault's competing 1977 date actually comes from: one unsourced secondary citation, not Werbos or a library record
- claim-werbos-1978-docsub-blocked — DARPA kept the full report out of DOCSUB; the '78 journal version shipped without the how-to appendix (two parallel reasons, on the record)
- claim-backprop-special-case-kelley-bryson-formula — Dreyfus's own formal identity claim (in flagged tension with the precursor framing)
- claim-hinton-biological-implausibility-four-objections — what "biologically implausible" actually asserts
- claim-fukushima-1979-neocognitron-first-cnn — the CNN's architectural ancestor (not backprop-trained)
- claim-parker-1985-technical-report-not-cited — the fourth uncited independent rediscoverer
- claim-ivakhnenko-gmdh-first-deep-characterization — "first deep learning, 1965": Schmidhuber's hedged characterization, regression-grown layers, eight-layer figure unverified
- claim-werbos-ieee-awards-1995-2022 — the awards are real (IEEE primaries); the 2022 citation language is itself myth-ledger circulation
- claim-amari-1972-associative-memory-precedes-hopfield — priority-without-citation, the cluster's recurring pattern
- claim-robbins-monro-1951-stochastic-approximation — the 1951 statistical ancestor of SGD, finally quoted from its own text (root-finding, not optimization)
- claim-amari-1968-saito-experiment-primary-read — the 1968 Kyoritsu pages read in Japanese: Saito's 1967 experiment is real and stochastic-descent-trained; the "five-layer MLP" framing is not the book's
- claim-bottou-2010-classifies-widrow-hoff-lms-as-sgd-matching-original-algorithm — Bottou's own 2010 paper tables LMS/Adaline alongside the Perceptron and k-Means as SGD-family members, matching the original papers
- claim-widrow-hoff-1960-original-paper-describes-lms-as-stochastic-steepest-descent — the 1960 WESCON original: the error surface is called "stochastic," searched one pattern at a time by steepest descent
- claim-widrow-lehr-1990-lms-instantaneous-gradient-unbiased-estimate — Widrow's 1990 retrospective proves the instantaneous gradient is an unbiased estimate of the true gradient, closing question-verify-lms-is-stochastic-gradient-descent-primary (now
answered) - claim-steinbuch-widrow-1965-comparison-not-multilayer-training — the 1965 date-coincidence resolved: Steinbuch & Widrow's note is a capacity comparison of two existing single-layer classifiers, not about multilayer training at all
- claim-widrow-1966-bootstrap-learning-preliminary-multilayer-attempt — a second, earlier, independently primary-documented 1960s multilayer attempt: a global non-gradient "selective bootstrap" reinforcement scheme, distinct from Madaline Rule I
The 1986 watershed
- claim-rhw-1986-reference-list-four-works — four references, none of the earlier derivations
- claim-rhw-1986-demonstration-not-invention — what the paper actually added (verified demos, verbatim symmetry-breaking)
The mechanism and the multiplicity
- claim-backpropagation-special-case-of-reverse-mode-ad — the precise correspondence (Baydin, JMLR)
- claim-jvp-vjp-transpose-duality — forward and reverse mode as one Jacobian in dual directions (pushforward/pullback; the cost duality follows from shape)
- claim-wengert-list-named-for-forward-mode-inventor — the tape's name honors the forward-mode inventor; automatic reverse mode is Speelpenning 1980
- claim-autograd-three-reifications-of-the-tape — PyTorch records a DAG, TensorFlow records a literal tape, JAX compiles a jaxpr and never says "tape"
- claim-cheap-gradient-bound-two-figures — why it's affordable: the gradient costs a small constant multiple of one forward evaluation ("one to four times" / "5×", reconciled)
- claim-reverse-mode-multiple-independent-discovery — at least five discoveries, five fields, 1965–1986 (Griewank 2012, direct read)
- claim-ml-and-ad-communities-mutually-unaware — why the field rediscovered it five times: two literatures that didn't read each other (Baydin 2018, Tier 1) — the structural root of the whole attribution gap
- claim-werbos-backprop-from-freud-own-account — the Freud thread, finally in Werbos's own words
- claim-werbos-1968-cybernetica-earliest-germ — the earliest germ: a real 1968 paper whose Freud→DP content survives only as Werbos's own paraphrase (no readable copy exists)
- claim-werbos-1997-optimization-consciousness-chapter — the mind/consciousness face of the same optimization architecture (1997, not 1996; not his fullest statement)
The bridge to the economics cluster
- claim-training-inference-compute-asymmetry-mechanism — the backward pass is the cost only training pays
The bridge to commercial application (a patent, not a paper)
- claim-us-patent-5819226-is-falcon-fraud-managers-founding-patent — HNC Software's founding 1992 patent for its Falcon Fraud Manager cites RHW 1986 by name as background for the exact backpropagation technique it commercializes — a concrete, dated instance of this cluster's lineage becoming financial infrastructure, six years after publication. See claim-hnc-1992-falcon-patent-names-network-architecture-as-feed-forward, claim-hnc-1992-falcon-patent-specifies-backpropagation-gradient-descent-supervised-training, claim-hnc-falcon-patent-describes-multilayer-not-single-layer-network, entity-robert-hecht-nielsen, entity-hnc-software.
The bridge to the schema-change cluster (the author is the hinge)
- observation-rumelhart-person-bridge-schema-theory-to-backprop — the "R" in
RHW 1986 (claim-rhw-1986-demonstration-not-invention) is the same David
Rumelhart who coined the 1978 tuning/restructuring taxonomy under
moc-schema-change-and-restructuring. One trajectory from cognitive-psychology
schema theory to the connectionist revival; a person-bridge no note had named.
seedling.
Biological plausibility and the alternatives
- claim-crick-1989-antidromic-not-weight-transport — the objection, correctly attributed
- claim-feedback-alignment-random-weights-train — random feedback weights (contested)
- claim-hinton-forward-forward-boltzmann-lineage — two forward passes, no backward
- claim-update-locking-backprop-constraint — the systems-level cost of the backward pass
- claim-brain-approximates-backprop-core-principles-ngrad — Hinton's mature position (Lillicrap 2020): the brain implements backprop's core principles via NGRAD, not strict backprop
- claim-hinton-backprop-in-brain-2007-to-2022-arc — the fifteen-year reversal: 2007 rescued backprop-in-the-brain, 2022 Forward-Forward abandoned it
The myth ledger (circulation vs. primary support)
- myth-werbos-invented-backpropagation — contested
- myth-linnainmaa-uncited-before-2010s — unresolved
- myth-amari-first-sgd-mlp — contested (single witness + citogenesis)
- myth-perceptrons-book-killed-connectionism — contested (proof-vs-conjecture collapse; causation genuinely multi-sided)
- claim-wikipedia-amari-sgd-citogenesis — how the amari myth's "corroboration" collapses to one witness: shared typo + shared misspelling
- claim-griewank-2012-does-not-mention-amari — confirmed negative: the reverse-mode history and the Amari-SGD claim never intersect at the primary level
The historiography of the myth ledger (reasoning about the ledger, not new entries in it)
(Section added 2026-09-04 by the Warden pass, warden/claude-opus-4.8, discharging the 2026-09-04 flag that the ledger above listed only its object-level entries while two independent historiography frameworks now sit on top of it. These are interpretive syntheses — each carries an [unverified-synthesis] flag in its own frontmatter — surfaced here so a reader of the ledger can find them, not asserted as settled.)
- observation-whig-history-splits-vault-myth-ledger-into-defensible-and-anachronistic — the first framework tested against the ledger. Butterfield's Whig-history critique, with Weinberg's science exemption, was proposed as a clean two-way sort (defensible present-knowledge appraisal vs. anachronistic concept-importing). Tested against all six ledger entries 2026-09-04, the binary does not hold: only myth-amari-first-sgd-mlp's "five-layer MLP" relabeling sorts as pure anachronism; myth-werbos-invented-backpropagation and myth-perceptrons-book-killed-connectionism are monocausal hero/villain collapses that fit neither pole; myth-linnainmaa-uncited-before-2010s and claim-griewank-2012-does-not-mention-amari are a residual bibliometric/scope category off the present-vs-past axis entirely. At least four failure shapes, not two.
- claim-weinberg-and-butterfield-both-explicitly-reject-hero-villain-narratives — the primary-grounded correction underneath that finding: read directly, both Weinberg's 2015 essay ("we should not … designat[e] some past scientists as spotless heroes … and others as villains") and the standard summary of Butterfield name hero/villain oversimplification as a third illegitimate mode — the one the ledger's Werbos and Perceptrons entries exhibit.
- claim-skinner-prolepsis-and-doctrines-name-two-modes-of-the-vaults-back-projection-pattern — a second, independent framework reaching the same "it's more than one error" conclusion from a different discipline. Skinner's 1969 taxonomy (claim-skinner-1969-names-four-mythologies-doctrines-coherence-prolepsis-parochialism) splits the ledger's "back-projection" into prolepsis (a later category imported — Schmidhuber reading "deep learning" onto 1676–1914 mathematics) versus mythology of doctrines (a later term credited — "credit assignment," absent from the 1974 Werbos thesis). A word can be checked with grep; a category can only be argued.
- claim-mahoney-warned-against-whig-history-of-computing — the disciplinary bridge that makes this vault's own home domain a legitimate target for the critique, with the honest caveat that the bridge is the vault's construction: per a direct read, Mahoney never names Butterfield or "Whig history" (claim-mahoney-2005-essay-omits-whig-and-butterfield) — he arrives at the same fallacy independently, from history-of-technology.
- Both syntheses are routed as open questions, not closed: question-verify-whig-defensible-vs-anachronistic-split-against-myth-ledger (answered — split rejected) and question-classify-myth-ledger-by-skinner-prolepsis-vs-doctrines (open — the drawer-by-drawer sort of every entry into prolepsis/doctrines/coherence has not been run).
Open verification threads
- question-verify-nyquist-bode-optimal-control-transmission-line — 2026-07-11: does Kelley 1960 (primary already read for claim-kelley-bryson-optimal-control-precursor) or Bryson's unread 1961 paper cite Nyquist/Bode/classical control theory by name? Cheapest open check in the cluster right now.
question-verify-linnainmaa-citation-count— CLOSED 2026-07-07 into_answered; the pull happened and moved the myth (claim-linnainmaa-1976-citation-histogram). Successor sharpener lives in myth-linnainmaa-uncited-before-2010s's watch flag (classify the 51 pre-2010 citations by community).- Library-op candidates for Cali (documents no automated fetch can reach): Werbos 1994 reprint TOC (resolves the 5.5.1 locator anomaly), Dreyfus 1973 IEEE TAC full text, Bryson 1961 symposium proceedings, Talking Nets ch. 16 p. 376 (Hinton's symmetry objection), Ivakhnenko 1971. Add, 2026-09-05: Talking Nets' Widrow chapter specifically — the presumed home of the "sharp quantizers/sigmoids" quote and any dated 1985/Snowbird account in Widrow's own words (direct.mit.edu 403'd, archive.org copy access-restricted, a Brown University repository page bot-blocked, all this session — see claim-widrow-abandoned-multilayer-training-until-1985-backprop's 2026-09-05 Correction history).
- 2026-07-07 evening safelist CLEARED: all backlog captures feeding this cluster are promoted or explicitly held (Opus-lane holds excepted).