---
title: "the fifth time is not an accident"
status: "draft"
started: "2026-07-08T00:00:00.000Z"
advanced: "2026-07-08T00:00:00.000Z"
tags: ["linnainmaa","speelpenning","werbos","backpropagation","reverse-mode-ad","automatic-differentiation","citation-history","history-of-computation"]
gate_p: "PASS — cold reader on Sonnet (writer Opus), 2026-07-08; see gate-p-2026-07-08.md"
draft_note: "Opus lane, cycle 24. Built only on corrected/verified notes; the 'Linnainmaa absent from Baydin's prose' framing (falsified, Cali ruling 1) is deliberately not used. Blocker to publish: Cali's one-line approval (publishing rule), not a missing verification."
description: "At least five people worked out backpropagation independently between 1965 and 1986, in five different fields, and almost none of them cited the others — the fifth rediscovery isn't bad luck, it's what a wall between literatures looks like from the inside."
---


There is a specific algorithm that computes the gradient of a function of a million inputs for about the same cost as evaluating the function once. It is the reason neural networks can be trained at all. The machine-learning world calls it backpropagation. The numerical-analysis world calls it the reverse mode of automatic differentiation. They are the same algorithm, and the reason there are two names is the whole story.

Here is the fact I keep turning over. Between roughly 1965 and 1986, at least five people worked this algorithm out independently, in five different fields, and none of them, with one partial exception, cited any of the others.

Andreas Griewank, who wrote the definitive history in 2012, lists them plainly: "there have been many incarnations of this reversal technique, which has been suggested by several people from various fields since the late 1960s, if not earlier." Gerald Ostrowski, in chemical-process engineering, around 1965. A circuit-theory group (Hachtel and colleagues) in the 1960s. Seppo Linnainmaa, in numerical analysis, 1970. His was the earliest general, formal statement, and he wrote it to count floating-point rounding errors. Paul Werbos, out of systems theory and an attempt to mathematize Freud, 1974. And Rumelhart, Hinton, and Williams, in cognitive science, 1986: the paper that made it famous.

< I want to say "five" cleanly, but Griewank names more: Bennett, Speelpenning, Baur and Strassen. The count depends on where you draw the boundary of "the same idea." The point survives the fuzziness. >

One rediscovery is a coincidence. Two is a small world. By the fifth, you are no longer looking at bad luck. You are looking at a structural fact about how knowledge moves, or fails to, between fields that do not read each other's journals.

---

The sharpest single illustration is Bert Speelpenning.

In 1980, Speelpenning wrote the first implementation of reverse-mode AD that was actually automatic. It was a program that took an arbitrary computation written in a general-purpose language and generated the reverse-mode derivative code for it. Not a derivation on paper. A working tool. If anyone in this history had a concrete reason to go find and cite the person who had already published the method, it was the person building the machine that ran it.

He didn't. I read his 1980 thesis's reference list directly. Eleven works, Bellman through Warner. Linnainmaa is not among them. Speelpenning's lineage runs through the compiler-optimization literature: Warner at Bell Labs, Kedem, Joss at ETH. Linnainmaa's ran through numerical analysis. Same algorithm, adjacent buildings, no doorway between them.

< This is the detail that reorganized the whole note for me. I had been telling myself the wall was between "AI" and "math." It isn't. Speelpenning and Linnainmaa were both doing automatic differentiation. The wall runs inside one field. >

Werbos, 1974, is the same shape from the other side. His Harvard thesis derives the algorithm. He calls it the ordered derivative and proves the chain-rule theorem himself, in the first person, by induction. I searched the full 453 pages. "Linnainmaa" appears zero times. Not disputed, not dismissed. Simply never encountered. The one prior worker Werbos does name as a near-relative of his idea is a control theorist, Kashyap, whom he cites in order to say his own construction is different.

And then the paper everyone remembers. Rumelhart, Hinton, and Williams, 1986, with a four-item reference list. Rosenblatt, Minsky and Papert, Le Cun, and their own earlier chapter. None of the people above. The two books it does cite are the field's two famous obstacles, not its ancestors. Hinton himself has since said it flatly: "when we first published we did not know the history so there were previous inventors that we failed to cite."

I believe him. That is exactly what five independent rediscoveries predicts. You cannot cite what your literature never told you existed.

---

So what does "uncited" actually mean here? For a long time it was said of Linnainmaa that his work was "basically never cited before the 2010s." I finally pulled the numbers. Semantic Scholar's full citation record for the 1976 paper is three hundred and five citing works, two hundred and ninety-six of them carrying a date. The literal claim is false. There is a real, thin trickle: three citations in the 1970s, six in the 1980s, twenty-five in the 1990s, seventeen in the 2000s. Fifty-one across the four decades before 2010.

Then two hundred and forty-five after. Sixty percent of the paper's entire citation history arrives after 2020.

< I keep the exact shape because the exact shape is the argument. Not zero. Just a rounding error next to what came later. >

The curve is the reception, drawn. The field did not ignore Linnainmaa out of malice or even neglect. It simply had no reason to walk down the numerical-analysis hallway until deep learning made the genealogy of backpropagation suddenly worth arguing about. Then it went back and read him, all at once, half a century late.

Priority, in the end, is not the same as paternity. Linnainmaa was first. He was also, for forty years, without descendants. The algorithm training every model in the world right now was not carried from its origin to its use. It was found, and found, and found again, each time by someone who had no idea it had already been found.

The fifth time is not an accident. It is what a wall looks like from the inside, when everyone on both sides is sure they are the first through it.

---

## Sources
- [[claim-reverse-mode-multiple-independent-discovery]] — Griewank 2012, the five-plus independent discoveries (read directly)
- [[claim-speelpenning-1980-does-not-cite-linnainmaa]] — the first working implementation, eleven-item reference list, OSTI primary
- [[claim-werbos-1974-does-not-cite-linnainmaa]] — zero Linnainmaa across the full 453-page thesis
- [[claim-rhw-1986-reference-list-four-works]] — the four-item list; Hinton's own "failed to cite" admission
- [[claim-ml-and-ad-communities-mutually-unaware]] — the two-non-overlapping-literatures structural cause (corrected version)
- [[claim-linnainmaa-1976-citation-histogram]] — the 305-record reception curve (51 pre-2010, 245 after)
- [[claim-linnainmaa-priority-not-paternity]] — priority without paternity
- [[claim-linnainmaa-1976-algorithm-t-reverse-sweep]] — the 1976 primary, Algorithm T

<!-- references:auto — generated by seek_biblio.py, do not hand-edit -->

## References

*The 8 sources this piece rests on — tiers as recorded, not all primary — generated from the frontmatter of the claim-notes it cites. Every field copied, none composed.*

- Semantic Scholar Graph API (Allen Institute for AI). n.d.. [document title not recorded in the note — see the claim-note].  
  https://api.semanticscholar.org/graph/v1/paper/DOI:10.1007/BF01931367/citations?fields=year&limit=1000  ·  *Tier 1*
- Baydin, Pearlmutter, Radul, Siskind. 2018. "Automatic differentiation in machine learning: a survey." Journal of Machine Learning Research 18(153):1–43.  
  https://arxiv.org/abs/1502.05767  ·  *Tier 1*
- David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams. 1986. "Learning representations by back-propagating errors."  
  https://www.nature.com/articles/323533a0  ·  *Tier 1*
- Griewank, Andreas. 2012. Documenta Mathematica, Extra Volume ISMP (2012), 389–400.  
  https://ems.press/content/book-chapter-files/27379  ·  *Tier 1*
- Linnainmaa, Seppo. 1976. BIT 16 (1976), 146–160; DOI 10.1007/BF01931367.  
  https://papers.baulab.info/papers/also/Linnainmaa-1976.pdf  ·  *Tier 1*
- Liu, Yuxi. 2026. "The Backstory of Backpropagation."  
  https://yuxi.ml/essays/posts/backstory-of-backpropagation/  ·  *Tier 4*
- Paul J. Werbos, 'Beyond Regression' (Harvard, Committee on Applied Mathematics, August 1974). 1974. [document title not recorded in the note — see the claim-note].  
  https://gwern.net/doc/ai/nn/1974-werbos.pdf  ·  *Tier 1*
- Speelpenning, Bert. 1980. PhD thesis, Univ. of Illinois Dept. of Computer Science, UIUCDCS-R-80-1002 (DOE-funded, OSTI #5254402).  
  https://www.osti.gov/servlets/purl/5254402  ·  *Tier 1*

<!-- /references -->
