---
title: "Did Shun'ichi Amari describe a form of gradient descent for layered networks in the 1960s?"
type: "capture"
status: "promoted"
date_promoted: "2026-07-06T00:00:00.000Z"
promoted_to: ["30-notes/myth-amari-first-sgd-mlp.md"]
not_promoted: ["All four claims: the load-bearing mechanism claim is single-witness (Schmidhuber) with the primary abstract pointing elsewhere — per RUN-QUEEN-LOOP Q5 this is myth-ledger material (`contested`), not claim-note material; the 07-03 follow-up captures' citogenesis evidence is folded into the myth entry","Claim (Amari's method was general SGD, not reverse-mode/chain-rule) — sound distinction but rests on the same single-witness chain; revisit if the 1968 book is ever read"]
origin: "batch"
date_created: "2026-06-29T00:00:00.000Z"
provenance: "batch run 2026-06-29; web research via WebSearch + WebFetch"
tags: ["amari","backpropagation","gradient-descent","stochastic-gradient-descent","multilayer-perceptron","history-of-ai","history-of-computation","japan","linnainmaa"]
primary_sources_consulted: [{"url":"https://arxiv.org/abs/2212.11279","author":"Jürgen Schmidhuber","date":"2022-12-21 (rev. 2025-12-29)","tier":1,"note":"'Annotated History of Modern AI and Deep Learning,' arXiv:2212.11279 [cs.NE]. Preprint/survey, author's own historical scholarship, not peer-reviewed in the conventional sense — treated as Tier 1 per the arXiv-clears-the-floor rule, same treatment as the Raugel et al. preprint in backpropagation-gap.md."},{"url":"https://people.idsia.ch/~juergen/deep-learning-history.html","author":"Jürgen Schmidhuber","date":"retrieved 2026-06-29 (updated periodically; mirrors the arXiv text)","tier":2},{"url":"https://people.idsia.ch/~juergen/who-invented-backpropagation.html","author":"Jürgen Schmidhuber","date":"retrieved 2026-06-29 (updated periodically)","tier":2},{"url":"https://en.wikipedia.org/wiki/Shun%27ichi_Amari","author":"Wikipedia contributors","date":"retrieved 2026-06-29","tier":3},{"url":"https://en.wikipedia.org/wiki/History_of_artificial_neural_networks","author":"Wikipedia contributors","date":"retrieved 2026-06-29","tier":3},{"url":"https://en.wikipedia.org/wiki/Multilayer_perceptron","author":"Wikipedia contributors","date":"retrieved 2026-06-29","tier":3},{"url":"https://theconversation.com/japanese-scientists-were-pioneers-of-ai-yet-theyre-being-written-out-of-its-history-243762","author":"Hansun Hsiung (Assistant Professor, Durham University)","date":"2024-11-27","tier":3},{"url":"https://people.idsia.ch/~juergen/amari1967.pdf","author":"Shun'ichi Amari","date":"1967","tier":1,"note":"Primary source located and fetched (1.5MB scanned PDF of IEEE Trans. Electronic Computers EC-16(3):299-307). Could NOT be OCR'd or quoted verbatim in this session — see access_failures. Bibliographic facts (title/journal/pages/DOI) corroborated independently via DBLP; substantive content claims about the paper rest on secondary characterization only."},{"url":"https://dblp.uni-trier.de/rec/journals/tc/Amari67.html","author":"DBLP","date":"retrieved 2026-06-29","tier":3,"note":"Bibliographic metadata only: title, journal, vol. 16, issue 3, pp. 299-307, DOI 10.1109/PGEC.1967.264666."}]
access_failures: ["https://www.durham.ac.uk/research/current/thought-leadership/2024/11/japanese-scientists-were-pioneers-of-ai-yet-theyre-being-written-out-of-its-history/ → 403 Forbidden (used The Conversation syndication of the same article instead)","https://grokipedia.com/page/Shun'ichi_Amari → 403 Forbidden","https://ieeexplore.ieee.org/document/1671260 → empty/blocked render","people.idsia.ch/~juergen/amari1967.pdf → fetched successfully (1.5MB) but is a scanned image PDF; no pdftotext/pdftoppm/poppler available in this sandbox to OCR it, so the primary text itself could not be directly quoted this session"]
---


**Short answer the claims below support:** Yes, with an important precision. Multiple convergent secondary sources (Schmidhuber's historical scholarship, three independent Wikipedia articles, a named-academic press piece) credit Amari's 1967 paper "A Theory of Adaptive Pattern Classifiers" with proposing the training of multilayer (layered) neural networks via stochastic gradient descent — and credit his student Saito with the first computer demonstration of this, published in Amari's 1968 Japanese-language book. This predates Rumelhart, Hinton & Williams (1986) by roughly two decades. But the historical record, read carefully, supports "a form of gradient descent for layered networks," not "Amari invented backpropagation": Schmidhuber's own account explicitly distinguishes Amari's general SGD method from the specific reverse-mode/chain-rule algorithm (the one [[claim-linnainmaa-reverse-mode-single-pass|Linnainmaa]] derived in 1970 and that became known as backpropagation), and states backpropagation is "generally more efficient." The primary 1967 paper itself was located but could not be directly read in this session (see access_failures), so the technical-mechanism claims below rest on Tier 1–2 secondary historiography rather than a direct reading of Amari's own words.

---

## Claim: Amari's 1967 paper is credited as the first proposal to train multilayer networks end-to-end by stochastic gradient descent

**Claim type:** Technical-mechanism / historical (load-bearing — "first" is a strong priority claim).

Jürgen Schmidhuber's historical survey states that in 1967 Amari "suggested to train MLPs with many layers in non-incremental end-to-end fashion from scratch by stochastic gradient descent (SGD), a method proposed in 1951 by Robbins & Monro." Schmidhuber's page elsewhere characterizes this as probably the first paper proposing SGD for learning in multilayer neural networks. Wikipedia's article on Amari independently states: "In 1967, he proposed the first deep learning artificial neural network (ANN) using the stochastic gradient descent (SGD) algorithm," citing Schmidhuber's arXiv survey as its source. The Wikipedia article on the history of artificial neural networks converges on the same claim: "The first deep learning multilayer perceptron trained by stochastic gradient descent was published in 1967 by Shun'ichi Amari."

**Sourcing floor check:** This is a technical-mechanism claim about what the paper proposed, and a historical-priority claim ("first"). It clears the floor via the Schmidhuber arXiv survey (Tier 1) and is corroborated by three independent Wikipedia pages (Tier 3, but all converge on the same Tier-1-sourced claim rather than being independent evidence). The underlying primary paper (Amari 1967, IEEE Trans. Electronic Computers, EC-16(3):299-307, DOI 10.1109/PGEC.1967.264666) was located and fetched but is a scanned PDF that could not be OCR'd in this session — its own wording on "multilayer" training was not directly verified. Note also a real tension: a search-engine-derived abstract summary of the 1967 paper (Tier 4, not independently verified) describes it as being about "error-correction adjustment procedures for determining the weight vector of linear pattern classifiers" and "piecewise-linear discriminant functions," with no explicit mention of multilayer networks — suggesting the "multilayer" framing credited to "1967" by historians may rely partly on Amari's related 1968 book (see next claim) rather than being fully explicit in the 1967 IEEE paper's own abstract. This tension is unresolved pending direct access to the primary text.

| Field | Value |
|---|---|
| source_url | https://arxiv.org/abs/2212.11279 (mirrored at https://people.idsia.ch/~juergen/deep-learning-history.html) |
| source_author | Jürgen Schmidhuber |
| source_date | 2022 (rev. 2025) |
| source_tier | 1 |
| exact_quote | "In 1967, however, Shun-Ichi Amari suggested to train MLPs with many layers in non-incremental end-to-end fashion from scratch by stochastic gradient descent (SGD), a method proposed in 1951 by Robbins & Monro." |
| corroborating_url | https://en.wikipedia.org/wiki/Shun%27ichi_Amari |
| corroborating_tier | 3 |
| corroborating_quote | "In 1967, he proposed the first deep learning artificial neural network (ANN) using the stochastic gradient descent (SGD) algorithm." |

---

## Claim: The empirical demonstration — a five-layer MLP with two modifiable layers — was Saito's, published in Amari's 1968 Japanese-language book, not in the 1967 English paper

**Claim type:** Technical-mechanism / historical.

Schmidhuber's account specifies that "Amari's implementation (with his student Saito) learned internal representations in a five layer MLP with two modifiable layers, which was trained to classify non-linearily separable pattern classes." The bibliographic record on Schmidhuber's site attributes this implementation to a separate publication from the 1967 IEEE paper: Amari, S. I. (1968), *Information Theory — Geometric Theory of Information*, Kyoritsu Publishing (in Japanese) — distinct from "A theory of adaptive pattern classifier," IEEE Trans. EC-16 (1967). Wikipedia's history-of-ANN article corroborates the experimental description independently: "According to Amari, in computer experiments conducted by his student Saito, a five layer MLP with two modifiable layers learned internal representations to classify non-linearly separable pattern classes," and the Multilayer Perceptron Wikipedia article gives the same figures: "a five-layered feedforward network with two learning layers."

Note a minor dating looseness across sources: Wikipedia's Amari article says Saito's report came "the same year" as the 1967 paper, while Schmidhuber's bibliographic citation dates the publication vehicle (the book) to 1968. The underlying experiment may have been conducted in 1967 and published in 1968; this has not been independently resolved.

**Sourcing floor check:** Technical-mechanism claim (what the experiment was, what architecture). Clears the floor via Schmidhuber (Tier 1/2), corroborated independently by two Wikipedia pages (Tier 3). The 1968 book itself (Japanese-language) was not located or read in this session — flag the specific architecture detail ("five layers, two modifiable layers") as resting on secondary historiography only, not on a primary reading. **[unverified-mechanism — needs primary: the 1968 book itself]**.

| Field | Value |
|---|---|
| source_url | https://people.idsia.ch/~juergen/who-invented-backpropagation.html |
| source_author | Jürgen Schmidhuber |
| source_date | retrieved 2026-06-29 |
| source_tier | 2 |
| exact_quote | "Amari's implementation (with his student Saito) learned internal representations in a five layer MLP with two modifiable layers, which was trained to classify non-linearily separable pattern classes." |
| primary_publication_cited | S. I. Amari (1968), Information Theory — Geometric Theory of Information, Kyoritsu Publ. (in Japanese) |
| corroborating_url | https://en.wikipedia.org/wiki/History_of_artificial_neural_networks |
| corroborating_tier | 3 |
| corroborating_quote | "According to Amari, in computer experiments conducted by his student Saito, a five layer MLP with two modifiable layers learned internal representations to classify non-linearly separable pattern classes." |

---

## Claim: Amari's method was a general stochastic gradient descent procedure, not the specific reverse-mode/chain-rule algorithm later called backpropagation

**Claim type:** Technical-mechanism (the precise distinction this whole question turns on).

Schmidhuber's history is explicit that Amari's approach should not be conflated with backpropagation proper. It frames Amari's contribution as a more general — but less computationally efficient — gradient method: "At least for supervised learning, backpropagation is generally more efficient than Amari's above-mentioned deep learning through the more general SGD method (1967), which learned useful internal representations in NNs about 2 decades earlier [than Rumelhart et al. 1986]." Elsewhere the same source describes the 1967 paper as using "stochastic gradient descent for learning in multilayer neural networks (without specifying the specific gradient descent method now known as reverse mode of automatic differentiation or backpropagation)."

This matters for the vault's existing backpropagation history: the efficient reverse-mode/chain-rule algorithm specifically is the one independently derived by [[claim-linnainmaa-reverse-mode-single-pass|Linnainmaa in 1970]] (see [[claim-linnainmaa-priority-not-paternity]] and [[backpropagation-gap]]) and later rediscovered by Werbos and by Rumelhart, Hinton & Williams. Amari's 1967 contribution is a parallel but distinct thread: a general (and reportedly less efficient) SGD procedure applied to a layered architecture, three years before Linnainmaa's specific reverse-mode formulation existed at all. The two threads are both "gradient descent for layered networks" in a loose sense, but they are not the same algorithm.

**Sourcing floor check:** Technical-mechanism claim distinguishing two algorithms. Clears the floor — both quotes come from Schmidhuber's survey, Tier 1 (arXiv version) / Tier 2 (mirrored site version), and the distinction is internally consistent across both versions of his text.

| Field | Value |
|---|---|
| source_url | https://people.idsia.ch/~juergen/deep-learning-history.html (text mirrors arXiv:2212.11279) |
| source_author | Jürgen Schmidhuber |
| source_date | retrieved 2026-06-29 |
| source_tier | 1 (arXiv version) / 2 (mirror) |
| exact_quote | "At least for supervised learning, backpropagation is generally more efficient than Amari's above-mentioned deep learning through the more general SGD method (1967), which learned useful internal representations in NNs about 2 decades earlier." |
| exact_quote_2 | "using stochastic gradient descent for learning in multilayer neural networks (without specifying the specific gradient descent method now known as reverse mode of automatic differentiation or backpropagation)" |

---

## Claim: Historians frame Amari's 1967 work as part of a pattern of early Japanese AI contributions sidelined from the dominant (Western, English-language) history of deep learning

**Claim type:** Historical / definitional (uncontested framing claim; no specific number attached).

Hansun Hsiung, an Assistant Professor at Durham University writing in The Conversation (November 2024, on the occasion of the 2024 Nobel Prize in Physics going to Hopfield and Hinton), frames Amari's 1967 work this way: "In 1967, Shun'ichi Amari proposed a method of adaptive pattern classification, which enables neural networks to self-adjust the way they categorise patterns, through exposure to repeated training examples." The same piece states that Amari's research "anticipated a similar method known as 'backpropagation,'" one of Hinton's credited contributions, and separately notes that Amari's 1972 associative-memory work was "mathematically equivalent" to Hopfield's 1982 paper, and that Kunihiko Fukushima built the first multilayer convolutional neural network in 1979. The article's overall thesis is that Japanese AI pioneers' contributions, including Amari's, have been comparatively written out of the standard Western narrative of deep learning's history.

This is structurally the same pattern already documented in the vault for [[claim-linnainmaa-priority-not-paternity|Linnainmaa]]: an early, mathematically substantive contribution to gradient-based learning, published in a non-English-dominant venue, that had little to no traceable causal influence on the AI community's later, independent arrival at the same territory. Unlike Linnainmaa's case, no specific claim has surfaced in this research session that Amari's work was completely uncited or unknown within Japan — only that it had limited influence on the international (especially Anglophone) AI research lineage that produced Rumelhart, Hinton & Williams (1986).

**Sourcing floor check:** Historical/definitional framing claim, uncontested in the sense that no source disputes Amari did this work — Tier 3 is acceptable per the rubric. (The Durham University press page itself returned 403; the identical content was confirmed via The Conversation, the original syndication source, with a named academic author and affiliation, which is the higher-credibility version of this claim anyway.)

| Field | Value |
|---|---|
| source_url | https://theconversation.com/japanese-scientists-were-pioneers-of-ai-yet-theyre-being-written-out-of-its-history-243762 |
| source_author | Hansun Hsiung (Assistant Professor, Durham University) |
| source_date | 2024-11-27 |
| source_tier | 3 |
| exact_quote | "In 1967, Shun'ichi Amari proposed a method of adaptive pattern classification, which enables neural networks to self-adjust the way they categorise patterns, through exposure to repeated training examples." |

---

## Further leads

- Amari's 1972 paper on associative memory is described (Tier 3, The Conversation) as "mathematically equivalent" to Hopfield's 1982 paper — worth checking against a primary source given Hopfield's 2024 Nobel Prize is the news hook for this framing.
- Kunihiko Fukushima built "the world's first multilayer convolutional neural network" in 1979, per the same Tier 3 source — a separate Japanese-AI-history thread, not yet checked against primary sources.
- The Robbins & Monro (1951) origin of stochastic gradient descent itself, cited in passing by Schmidhuber as the method Amari applied to MLPs in 1967 — not independently verified this session.
- Primary text of Amari (1967), "A Theory of Adaptive Pattern Classifiers," IEEE Trans. Electronic Computers EC-16(3):299-307 — PDF located at https://people.idsia.ch/~juergen/amari1967.pdf but is a scanned image; needs OCR tooling (poppler/pdftotext, unavailable in this sandbox) or a non-scanned copy to directly quote.
- Amari's 1968 book *Information Theory — Geometric Theory of Information* (Kyoritsu, in Japanese) — the actual publication vehicle for Saito's five-layer MLP experiment; not located or read this session.
- Griewank (2012), "Who Invented the Reverse Mode of Differentiation?" (already a lead in the Linnainmaa capture chain — see [[claim-linnainmaa-reverse-mode-single-pass]]) — check whether it discusses Amari at all, since it is the dedicated scholarly priority paper for the adjacent Linnainmaa case.
- Schmidhuber's "Japanese scientists were pioneers" framing overlaps with his broader, long-running campaign to re-credit non-Anglophone originators of deep-learning techniques (also visible in his Linnainmaa advocacy) — worth noting as a possible source bias if this capture is promoted: Schmidhuber is simultaneously the best-documented Tier 1 source on Amari's priority and an advocate with a track record of pushing priority claims that benefit a particular historiographic narrative.

> [!note] Seek's commentary:
> The headline answer is yes, but the interesting finding is the seam between Claims 1 and 3: the popular shorthand ("Amari did backpropagation in 1967") that shows up in search-engine summaries is imprecise in exactly the way Schmidhuber's own primary text warns against — general SGD applied to a layered network is not the same claim as the specific reverse-mode chain-rule algorithm. That distinction is exactly the one [[backpropagation-gap]] already cares about for Linnainmaa, and it's worth holding the line on when this gets promoted: "gradient descent for layered networks" (true, well-sourced) is a different and weaker claim than "invented backpropagation" (not what the Tier 1 source actually says). The other live worry is that nearly every claim in this capture traces back to a single historian (Schmidhuber) as the ultimate Tier 1/2 source, even where Wikipedia appears to "corroborate" — those Wikipedia passages cite Schmidhuber directly, so they are not independent confirmation, just propagation. A promoted claim-note should say so plainly rather than letting three Wikipedia pages look like three sources.
