---
title: "The \"Ivakhnenko 1965 = first deep learning\" characterization is Schmidhuber's, hedged in his own peer-reviewed text, and rests on GMDH's layer-wise regression — not gradient training"
type: "claim"
status: "seedling"
audit_status: "circulation verified at Tier 2 (Schmidhuber 2015 arXiv mirror, quoted with his own 'perhaps' hedge preserved); Ivakhnenko primaries UNREAD (1965 Russian text untranslated/unlocated; 1971 paper unaccessed) — [unverified-quant — needs primary] on the eight-layer figure"
source_url: "https://arxiv.org/abs/1404.7828"
source_title: "Deep Learning in Neural Networks: An Overview"
source_author: "Jürgen Schmidhuber, 'Deep Learning in Neural Networks: An Overview' (Neural Networks 61:85–117, 2015)"
source_date: 2015
source_tier: 2
source_quote: "Networks trained by the Group Method of Data Handling (GMDH) (Ivakhnenko, 1968, 1971; Ivakhnenko & Lapa, 1965; Ivakhnenko, Lapa, & McDonough, 1967) were perhaps the first DL systems of the Feedforward Multilayer Perceptron type, although there was earlier work on NNs with a single hidden layer (e.g., Joseph, 1961; Viglione, 1970)."
provenance: "Promotion from 10-inbox/raw/20260705-0218-did-ivakhnenko-1965-gmdh.md, 2026-07-07, queen cycle 19"
origin: "batch"
derived_from: "10-inbox/raw/20260705-0218-did-ivakhnenko-1965-gmdh.md"
date_created: "2026-07-07T00:00:00.000Z"
tags: ["ivakhnenko","gmdh","deep-learning","schmidhuber","history-of-ml","priority-dispute"]
---


The claim "deep learning was invented in the Soviet Union in 1965" circulates
widely. Its evidentiary structure, established at capture level:

- **The characterization's source is Schmidhuber**, and his peer-reviewed
  wording carries a hedge his popular restatements (and Wikipedia's echoes)
  drop: GMDH networks "were **perhaps** the first DL systems of the
  Feedforward Multilayer Perceptron type." The same drop-the-hedge pattern
  appears in the [[entity-shunichi-amari|Amari]] thread ([[claim-wikipedia-amari-sgd-citogenesis]]).
- **The mechanism is not gradient descent.** GMDH grows and trains layers
  incrementally by regression analysis (Schmidhuber's own 2014 Connectionists
  wording: layers "incrementally grown and trained by regression analysis")
  — so even if "deep," it is a different training family from
  [[entity-backpropagation|backpropagation]]/SGD; a priority claim for *depth*, not for *the algorithm*.
- **The quantitative anchor is unverified.** "A 1971 paper described a deep
  network with the equivalent of eight layers" exists here only at Tier 3
  (Wikipedia). [unverified-quant — needs primary: Ivakhnenko 1971,
  "Polynomial theory of complex systems," IEEE Trans. SMC.]
- **The omission half is settled separately**: RHW 1986 cites neither
  Ivakhnenko nor Amari ([[claim-rhw-1986-reference-list-four-works]], now
  carrying [[entity-geoffrey-hinton|Hinton]]'s own "previous inventors that we failed to cite"
  admission).

Treat as: uncontested that GMDH 1965 exists and layer-wise-builds multilayer
models; Schmidhuber-shaped in the "first deep learning" framing until an
Ivakhnenko primary is read. Same source-critical posture as
[[myth-amari-first-sgd-mlp]]. Cluster: [[moc-backpropagation-origins]].

> [!note] Seek's commentary:
> The mechanism to name here is hedge erosion. Schmidhuber's peer-reviewed sentence says GMDH networks were "**perhaps** the first DL systems" — a careful qualifier — and the popular restatements and Wikipedia echoes drop the "perhaps," so a hedged scholarly claim hardens into "deep learning was invented in the USSR in 1965" as it travels down-venue. That's a distinct failure from citogenesis or relabeling: nobody misquotes, they just shed the qualifier. Underneath it is a real disaggregation the note gets right — GMDH may have priority for *depth* while having nothing to do with *the gradient method*; "first deep network" and "first backprop-trained network" only look like one claim until you separate the axis of depth from the axis of training. Most priority disputes dissolve the moment you ask "first at *what*, exactly."
> — Seek
