Fukushima's Neocognitron (1979/1980) introduced the core CNN architecture — weight-shared local features plus downsampling — but was not trained by backpropagation
The Neocognitron introduced what became the convolutional neural network's defining structure: layers of weight-shared local feature detectors alternating with downsampling ("S-cells" and "C-cells"), giving shift-invariant pattern recognition. Fukushima's own 1980 abstract frames it in terms of unsupervised self-organization — not "convolution," not "first," and with no priority claim. The "first CNN" credit is historians' (chiefly Schmidhuber, who calls it "perhaps the first artificial NN that deserved the attribute deep" — "Deep Learning in Neural Networks: An Overview", arXiv 1404.7828 §5.4, pointer added 2026-09-11 audit; the same paragraph supplies the qualifier: "Fukushima, however, did not set the weights by supervised backpropagation … but by local, WTA-based unsupervised learning rules"), with the sharp qualifier the retellings often drop: it was not trained by backpropagation — its learning was unsupervised self-organization, and the backprop-trained CNN is LeCun 1989 (claim-lecun-1989-first-practical-recognition).
So the honest form: Neocognitron is the architectural ancestor of the CNN, LeCun 1989 is where that architecture met backpropagation. Secondary sources disagree on how strongly to state the priority — even two Wikipedia articles use different strengths — which is the same "who invented it" compression the myth ledger tracks. The Schmidhuber single-witness caveat applies (entity-juergen-schmidhuber). See moc-backpropagation-origins.
Source
“Neocognitron”