---
id: "20260705-0218-did-ivakhnenko-1965-gmdh"
title: "Did Ivakhnenko (1965 GMDH) and Amari (1967 SGD MLP) develop deep-network training methods that Rumelhart-Hinton-Williams 1986 failed to cite?"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-ivakhnenko-gmdh-first-deep-characterization.md (new: hedge-preserving, mechanism-distinguishing, eight-layer figure flagged unverified-quant)","30-notes/claim-rhw-1986-reference-list-four-works.md (update section: Hinton's own failed-to-cite admission + verbatim four-entry list)"]
not_promoted: ["Amari half — superseded by cycle 18's primary read (claim-amari-1968-saito-experiment-primary-read); this capture's [unverified-mechanism] flags on the five-layer claim are now partially resolved there (experiment real, MLP framing absent)","Schmidhuber's pointed 'failed to cite the origin' phrasing — search-snippet-only, never confirmed against the live page; stays unsourced in the capture"]
promotion_note: "Queen cycle 19, 2026-07-07."
origin: "batch"
model: "claude-sonnet-5"
date_created: "2026-07-05T00:00:00.000Z"
provenance: "batch run 2026-07-05; web research via a background research subagent (WebSearch + WebFetch) plus direct follow-up verification by the main session, including a direct Read of the OCR'd text of the fetched Nature 1986 PDF; harvested from 2026-06-29-what-did-the-1986-rumelhart-hinton-williams-nature-paper-add-that-earlier-formulations-did-not"
derived_from: []
tags: ["backpropagation","rumelhart-hinton-williams-1986","ivakhnenko","gmdh","amari","stochastic-gradient-descent","history-of-ai","priority-disputes","primary-source-verification"]
primary_sources_consulted: [{"url":"https://gwern.net/doc/ai/nn/1986-rumelhart-2.pdf","author":"David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams","date":"1986-10-09","tier":1,"note":"Direct mirror/scan of the actual Nature letter, 'Learning representations by back-propagating errors,' Nature 323, pp. 533-536 (received 1 May, accepted 31 July 1986). A WebFetch summarization pass reported the PDF as unreadable/binary; the Read tool applied directly to the same saved PDF returned full OCR'd text of all four pages, including the complete reference list, which is quoted verbatim below."},{"url":"https://arxiv.org/abs/1404.7828","author":"Jürgen Schmidhuber","date":"2015 (Neural Networks 61:85-117; arXiv preprint mirror of the published, peer-reviewed version)","tier":2,"note":"'Deep Learning in Neural Networks: An Overview.' Peer-reviewed venue, but Schmidhuber is the central, long-running claimant in the deep-learning priority dispute this capture concerns — treat his characterizations as an interested party's, not a neutral historian's, even though the venue clears the technical-mechanism sourcing floor."},{"url":"https://en.wikipedia.org/wiki/History_of_artificial_neural_networks","author":"Wikipedia contributors","date":"retrieved 2026-07-05","tier":3,"note":"Used only for the specific numeric/mechanism details (eight-layer 1971 GMDH network; five-layer Amari/Saito SGD experiment) that could not be independently verified against primary or Tier 1-2 text in this session. Flagged accordingly below, not treated as clearing the floor on its own."},{"url":"https://en.wikipedia.org/wiki/Multilayer_perceptron","author":"Wikipedia contributors","date":"retrieved 2026-07-05","tier":3,"note":"Same caveat as above; this page's Amari/Saito claim is itself footnoted to Amari (1967) primary and to Schmidhuber's 2022 'Annotated History of Modern AI and Deep Learning' (arXiv:2212.11279), which could not be independently opened in this session."},{"url":"https://ieeexplore.ieee.org/document/4039068/","author":"Shun'ichi Amari","date":"1967","tier":1,"note":"'A Theory of Adaptive Pattern Classifiers,' IEEE Transactions on Electronic Computers EC-16(3):299-307. Confirmed to exist and resolve; only the abstract was retrievable (full text paywalled on IEEE Xplore). The retrievable abstract text discusses error-correction convergence for linear pattern classifiers and does not itself mention a multilayer/hidden-unit network."},{"url":"https://syncedreview.com/2020/04/23/who-invented-backpropagation-hinton-says-he-didnt-but-his-work-made-it-popular/","author":"Fangyu Cai and Yuan Yuan, Synced (SyncedReview)","date":"2020-04-23","tier":2,"note":"Journalism piece, but the load-bearing content is a direct quotation of Geoffrey Hinton's own public (Reddit) statement, independently verified by re-fetching this URL directly; treated as Tier 2 for the quoted admission itself, since it is the primary actor's own words relayed by a named, byline-attributed outlet, not an aggregator's paraphrase."},{"url":"https://mailman.srv.cs.cmu.edu/pipermail/connectionists/2014-July/027186.html","author":"Jürgen Schmidhuber, posted to the Connectionists mailing list","date":"2014-07","tier":3,"note":"Schmidhuber's own words, self-hosted on a public but self-authored channel; independently re-fetched and confirmed to resolve. Self-interested source for the priority claim; recorded as evidence that Schmidhuber has made this specific argument in his own words, not as neutral corroboration of the argument's substance."}]
access_failures: [{"url":"https://people.idsia.ch/~juergen/who-invented-backpropagation.html","reason":"DNS resolution failure (ENOTFOUND) on every attempt, both via the background research subagent and via direct follow-up in the main session. This is Schmidhuber's own primary essay making the 'failed to cite' argument at length; could not be read directly in this session. A search-engine snippet surfaced closely matching phrasing (see commentary), but this was not independently confirmed against the live page."},{"url":"https://people.idsia.ch/~juergen/amari1967.pdf","reason":"Same DNS failure; this is Schmidhuber's hosted copy of Amari's 1967 paper and could not be read to verify the 'five-layer, two-modifiable-layer, Saito experiment' claim against the primary text."},{"url":"https://people.idsia.ch/~juergen/physics-nobel-2024-plagiarism.pdf","reason":"Same DNS failure. Confirmed only via search-result titles to exist; this is Schmidhuber's most recent (Oct. 2024) restatement of the priority argument, in the context of the Hopfield/Hinton Nobel Prize in Physics."},{"url":"https://www.nature.com/articles/323533a0","reason":"Redirects to an authentication/paywall gate; publisher's own hosted full text not accessible without institutional login. Substituted with the gwern.net mirror, which was read directly and matches this record's own DOI/citation metadata (independently cross-checked against a Scientific Research Publishing reference page, https://www.scirp.org/reference/referencespapers?referenceid=1698775, which lists the identical authors/year/journal/volume/pages/DOI)."},{"url":"https://ieeemilestones.ethw.org/Milestone-Proposal:Invention_of_Deep_Learning,_1965","reason":"HTTP 403 Forbidden on direct fetch; surfaced by search as a possible additional corroborating source (an IEEE Milestones proposal specifically for Ivakhnenko & Lapa's 1965 work) but its content could not be read or quoted in this session."}]
---


**Short answer:** The citation-omission half of this question is directly confirmed — the Nature 1986 letter's complete, four-item reference list does not include Ivakhnenko, Amari, Werbos, or Linnainmaa, and Hinton has separately acknowledged in his own words that "there were previous inventors that we failed to cite." The substantive half — whether Ivakhnenko's 1965 GMDH and Amari's 1967 paper actually constitute deep-network training methods in a sense comparable to backpropagation — rests heavily on Jürgen Schmidhuber's own historical reconstruction (a peer-reviewed but self-interested source, since Schmidhuber is the central claimant in an ongoing priority dispute over deep learning credit), with only partial independent corroboration and no fully independent verification of the specific "eight-layer" (Ivakhnenko) and "five-layer, Saito experiment" (Amari) figures in this session.

---

## Claim: Rumelhart, Hinton & Williams (1986)'s complete reference list cites only four works, none of them Ivakhnenko or Amari

**Claim type:** Historical/bibliographic fact, directly checkable against the primary document. Because it is the load-bearing point of the whole question, it is held to the Tier 1-2 floor regardless.

The Nature letter's reference list, read directly and in full from the OCR'd primary-source scan, consists of exactly these four entries:

1. "Rosenblatt, F. *Principles of Neurodynamics* (Spartan, Washington, DC, 1961)."
2. "Minsky, M. L. & Papert, S. *Perceptrons* (MIT, Cambridge, 1969)."
3. "Le Cun, Y. *Proc. Cognitiva* 85, 599-604 (1985)."
4. "Rumelhart, D. E., Hinton, G. E. & Williams, R. J. in *Parallel Distributed Processing: Explorations in the Microstructure of Cognition. Vol. 1: Foundations* (eds Rumelhart, D. E. & McClelland, J. L.) 318-362 (MIT, Cambridge, 1986)."

The paper's only other acknowledgment of prior/independent discovery is a single in-text sentence, immediately following the momentum-update equation: "Variants on the learning procedure have been discovered independently by David Parker (personal communication) and by Yann Le Cun." No mention of Ivakhnenko, Amari, Werbos, or Linnainmaa appears anywhere in the four pages of the letter.

**Sourcing floor check:** Cleared — Tier 1, the primary letter itself, read directly.

| Field | Value |
|---|---|
| source_url | https://gwern.net/doc/ai/nn/1986-rumelhart-2.pdf |
| source_author | David E. Rumelhart, Geoffrey E. Hinton, Ronald J. Williams |
| source_date | 1986-10-09 |
| source_tier | 1 |
| exact_quote | Full reference list as transcribed above, plus: "Variants on the learning procedure have been discovered independently by David Parker (personal communication) and by Yann Le Cun." |

---

## Claim: Ivakhnenko & Lapa's 1965 GMDH work is characterized (by Schmidhuber) as "perhaps the first DL systems of the Feedforward Multilayer Perceptron type"

**Claim type:** Specific technical-mechanism / historical characterization — surprising and contested (it is a priority claim), so escalated to the Tier 1-2 floor per the rubric even though "who did what, when" claims can otherwise rest at Tier 3-4.

Schmidhuber's peer-reviewed 2015 survey states: "Networks trained by the Group Method of Data Handling (GMDH) (Ivakhnenko, 1968, 1971; Ivakhnenko & Lapa, 1965; Ivakhnenko, Lapa, & McDonough, 1967) were perhaps the first DL systems of the Feedforward Multilayer Perceptron type, although there was earlier work on NNs with a single hidden layer (e.g., Joseph, 1961; Viglione, 1970)." Note the hedge ("perhaps") in Schmidhuber's own wording.

This is echoed, without the hedge, by Wikipedia's "History of artificial neural networks" article: "Group method of data handling, a method to train arbitrarily deep neural networks, was published by Alexey Ivakhnenko and Valentin Lapa in 1965," which the article says they "regarded ... as a form of polynomial regression or a generalisation of Rosenblatt's perceptron." Wikipedia cites both Ivakhnenko's own primary papers and Schmidhuber as sources for this characterization; the primary Ivakhnenko/Lapa texts themselves were not accessed or read directly in this session (untranslated from Russian and not found in a fetchable full-text mirror).

**Sourcing floor check:** Partially cleared. The core characterization ("first DL/MLP-type system trained via GMDH") is sourced to a peer-reviewed Tier 2 venue (Schmidhuber 2015), acceptable per the floor, but flagged because Schmidhuber has a direct personal stake in this specific priority narrative — record the characterization as *his claim*, corroborated but not independently re-derived from Ivakhnenko's original text.

| Field | Value |
|---|---|
| source_url | https://arxiv.org/abs/1404.7828 (arXiv mirror of Schmidhuber, J. (2015), *Neural Networks* 61:85-117) |
| source_author | Jürgen Schmidhuber |
| source_date | 2015 |
| source_tier | 2 (peer-reviewed venue; interested-party caveat applies) |
| exact_quote | "Networks trained by the Group Method of Data Handling (GMDH) (Ivakhnenko, 1968, 1971; Ivakhnenko & Lapa, 1965; Ivakhnenko, Lapa, & McDonough, 1967) were perhaps the first DL systems of the Feedforward Multilayer Perceptron type, although there was earlier work on NNs with a single hidden layer (e.g., Joseph, 1961; Viglione, 1970)." |

### Sub-claim: a 1971 Ivakhnenko paper described an eight-layer network trained this way

This specific number is only available via Wikipedia in this session: "A 1971 paper described a deep network with the equivalent of eight layers trained by this method." This is a quantitative claim (a specific layer count) resting only on a Tier 3 source.

`[unverified-quant -- needs primary: confirm the "eight layers" figure directly against Ivakhnenko's 1971 IEEE Trans. SMC paper, "Polynomial theory of complex systems," which was not accessed in this session]`

---

## Claim: Amari (1967) trained a multilayer network by stochastic gradient descent

**Claim type:** Specific technical-mechanism claim (what the 1967 system actually did) plus, for the layer count, a quantitative claim. Requires Tier 1-2.

The primary paper — Amari, S. (1967), "A Theory of Adaptive Pattern Classifiers," *IEEE Transactions on Electronic Computers* EC-16(3):299-307 — was confirmed to exist (https://ieeexplore.ieee.org/document/4039068/), but only its abstract was retrievable in this session (full text paywalled). That abstract discusses error-correction convergence procedures for the weight vector of *linear* pattern classifiers under nonseparable distributions, and does **not** itself mention a multilayer or hidden-unit network.

The specific claim that this 1967 work included a five-layer MLP trained by SGD comes from Wikipedia's "Multilayer perceptron" article: "Shun'ichi Amari reported the first multilayered neural network trained by stochastic gradient descent, was able to classify non-linearly separable pattern classes. Amari's student Saito conducted the computer experiments, using a five-layered feedforward network with two learning layers." Wikipedia footnotes this to Amari (1967) directly and to Schmidhuber's 2022 "Annotated History of Modern AI and Deep Learning" (arXiv:2212.11279) — neither of which was independently opened and read in this session (Amari's full text is paywalled; the Schmidhuber 2022 arXiv paper was not fetched).

**Sourcing floor check:** NOT cleared at Tier 1-2 for the specific mechanism/layer-count detail — recorded and flagged per the sourcing floor rule.

`[unverified-mechanism -- needs primary: confirm against Amari's full 1967 text, or against Amari's 1968 Japanese-language book (Kyoritsu) reportedly describing the Saito experiment, that a five-layer network with two modifiable/learning layers was actually trained by SGD]`
`[unverified-quant -- needs primary: the "five layers, two learning layers" figure specifically]`

Note: a related capture in this vault, `2026-07-03-does-amaris-1968-japanese-book-kyoritsu-contain-saitos-five-layer-mlp-experiment.md`, appears to already be investigating this exact sub-question directly — cross-reference before promotion to avoid duplicating that research.

| Field | Value |
|---|---|
| source_url | https://en.wikipedia.org/wiki/Multilayer_perceptron |
| source_author | Wikipedia contributors |
| source_date | retrieved 2026-07-05 |
| source_tier | 3 |
| exact_quote | "Shun'ichi Amari reported the first multilayered neural network trained by stochastic gradient descent, was able to classify non-linearly separable pattern classes. Amari's student Saito conducted the computer experiments, using a five-layered feedforward network with two learning layers." |

---

## Claim: Hinton has publicly acknowledged that prior inventors went uncited in the original publication

**Claim type:** Historical/biographical — a direct admission by one of the paper's own authors, so treated as strong corroboration even though relayed via a journalism outlet rather than a primary transcript.

Synced (SyncedReview), in a 2020 piece by named journalists Fangyu Cai and Yuan Yuan, reports Hinton's own Reddit response to the priority dispute: "It is true that when we first published we did not know the history so there were previous inventors that we failed to cite." The article frames this as Hinton conceding an innocent gap in awareness of prior work, not (as Schmidhuber argues) a deliberate omission.

**Sourcing floor check:** Cleared for the fact that Hinton made this statement (verified by directly re-fetching the article and confirming the quotation) — treated as Tier 2 because the substance is the primary actor's own words, independently reported by a named byline, not an aggregator paraphrase. The underlying original Reddit post itself was not independently located in this session.

| Field | Value |
|---|---|
| source_url | https://syncedreview.com/2020/04/23/who-invented-backpropagation-hinton-says-he-didnt-but-his-work-made-it-popular/ |
| source_author | Fangyu Cai and Yuan Yuan (Synced), quoting Geoffrey Hinton |
| source_date | 2020-04-23 |
| source_tier | 2 |
| exact_quote | "It is true that when we first published we did not know the history so there were previous inventors that we failed to cite." |

---

## Claim: Schmidhuber has explicitly and repeatedly argued that RHW86 failed to cite Ivakhnenko and Amari specifically

**Claim type:** Historical claim about the existence and content of an argument — the argument's *existence* is uncontested and independently corroborated below; its *correctness* is the contested part addressed in the claims above.

Schmidhuber, posting to the public Connectionists mailing list in July 2014 (independently re-fetched and confirmed to resolve), wrote: "Compare also the first adaptive, deep, multilayer perceptrons (Ivakhnenko et al., since 1965), whose layers are incrementally grown and trained by regression analysis" and separately, of the 1986 paper: "A paper of 1986 significantly contributed to the popularisation of BP (Rumelhart et al., 1986)." This confirms, in Schmidhuber's own words on a durable public archive, that he has made substantially this argument.

A more direct and pointed statement — reportedly reading "[RHW86] did not cite the origin of the method, also known as the reverse mode of automatic differentiation... [and] failed to cite the first working algorithms for deep learning of internal representations (Ivakhnenko & Lapa, 1965) as well as Amari's work (1967-68)" — surfaced repeatedly across independent search-engine snippets, apparently drawn from Schmidhuber's essay "Who Invented Backpropagation?" at people.idsia.ch. That page itself could not be reached in this session (DNS resolution failure on every attempt, both via a background research subagent and via direct follow-up), so this more pointed wording is `[unsourced -- needs verification]` against the live primary text and is not treated as a confirmed direct quotation here — only the mailing-list quote above is recorded as verified.

| Field | Value |
|---|---|
| source_url | https://mailman.srv.cs.cmu.edu/pipermail/connectionists/2014-July/027186.html |
| source_author | Jürgen Schmidhuber |
| source_date | 2014-07 |
| source_tier | 3 (self-authored, self-interested priority claimant; public archive, independently re-verified to resolve) |
| exact_quote | "Compare also the first adaptive, deep, multilayer perceptrons (Ivakhnenko et al., since 1965), whose layers are incrementally grown and trained by regression analysis" / "A paper of 1986 significantly contributed to the popularisation of BP (Rumelhart et al., 1986)." |

---

## What remains unresolved

Two access failures materially limit this capture: (1) Schmidhuber's own primary essay making the detailed "failed to cite" argument (people.idsia.ch) was unreachable throughout this session, so the strongest, most explicit statement of the claim is not independently confirmed verbatim here — only paraphrased search-snippet text and a shorter, related 2014 mailing-list quote were verified; (2) the full texts of both Ivakhnenko & Lapa (1965) and Amari (1967) were not independently read (untranslated/paywalled respectively), so the specific technical claims about layer counts and what exactly was trained by what procedure rest on secondary characterization (chiefly Schmidhuber's, adopted by Wikipedia editors) rather than on this session's own reading of the primary texts.

> [!note] Seek's commentary:
> The two halves of this question have very different evidentiary footing. "RHW86 didn't cite Ivakhnenko or Amari" is now about as solid as a historical claim gets — read directly off the primary source's own four-item bibliography, with Hinton's own public statement as corroboration that the gap was real (if, per Hinton, unintentional). "Ivakhnenko's GMDH and Amari's 1967 paper were themselves deep-network training methods in a sense that made the omission a *meaningful* one" is a much softer claim that traces back overwhelmingly to one interested party's historical reconstruction. That doesn't make Schmidhuber wrong — GMDH's status as an early deep, layer-wise-trained system appears to be reasonably well accepted independent of him — but the specific comparative framing ("first," "deep," "SGD," exact layer counts) is doing a lot of work in his phrasing specifically, and this capture could not independently verify it against Ivakhnenko's or Amari's own texts. Treat the "did they fail to cite it" half as settled and the "was it deep-network training" half as Schmidhuber-shaped until a primary source closes the gap.

**Central question status:** Partially confirmed. The citation omission itself — RHW86 (1986) cites neither Ivakhnenko nor Amari, and Hinton has since acknowledged failing to cite "previous inventors" — is confirmed directly from primary and near-primary sources. Whether Ivakhnenko's 1965 GMDH and Amari's 1967 paper constitute "deep-network training methods" in a sense that makes this omission historically significant is `[unverified -- could not confirm or deny after search]` at the Tier 1-2 level required for a technical-mechanism claim of this kind; it currently rests on Schmidhuber's own (peer-reviewed but self-interested) characterization, echoed by Wikipedia, with the specific quantitative details (eight-layer GMDH network, five-layer Amari/Saito SGD experiment) unverified against primary texts.
