---
title: "Myth ledger: \"LeCun's 1988 paper used hand-designed convolutional kernels\" — the 1988 hand-designed-kernel ZIP-code paper was Denker et al., not LeCun"
type: "myth"
status: "budding"
circulating_claim: "A '1988 LeCun paper' used hand-designed (rather than learned) convolutional kernels, and that is why LeCun's 1989 work is called the first end-to-end trained convnet."
where_it_circulates: "Secondary retellings including Wikipedia's LeNet and CNN articles, and the vault's own earlier harvest lead — all shorthand 'LeCun 1988' for a paper LeCun did not author."
primary_source_status: "contested"
audit_status: "primary-verified (Denker et al. 1988 read directly via pdftotext 2026-07-09 after the poppler repair — 9pp, clean text). Two facts separated: (1) authorship — LeCun is NOT on the 1988 paper (DEBUNKED for the 'LeCun 1988' label); (2) hand-designed feature detectors — CONFIRMED in the primary text."
source_url: "https://proceedings.neurips.cc/paper/1988/file/a97da629b098b75c294dffdc3e463904-Paper.pdf"
source_author: "J. S. Denker, W. R. Gardner, H. P. Graf, D. Henderson, R. E. Howard, W. Hubbard, L. D. Jackel, H. S. Baird, I. Guyon (AT&T Bell Labs)"
source_date: 1988
source_tier: 1
source_quote: "Feature b is designed to detect the right-hand end of (approximately) horizontal strokes."
provenance: "Promotion from 10-inbox/raw/2026-07-09-did-lecuns-1988-paper-use-hand-designed-rather.md, 2026-07-09, Fable clean-lane promotion marathon; the capture's [unverified-mechanism] flag was cleared by a direct read of the Denker et al. 1988 NIPS paper."
origin: "batch"
derived_from: ["20260709-0233-did-lecuns-1988-paper"]
date_created: "2026-07-09T00:00:00.000Z"
tags: ["lecun","denker","bell-labs","convolutional-neural-networks","backpropagation","history-of-ml","zip-codes","priority","citation-practice"]
audits: ["2026-07-09 claude-fable-5"]
drafted_in: ["a-name-is-a-summary"]
---


**The circulating claim** compresses two Bell Labs ZIP-code papers into one
"LeCun 1988 → 1989" story. The 1988 paper actually invoked for hand-designed
kernels is **"Neural Network Recognizer for Hand-Written Zip Code Digits," by
J. S. Denker, W. R. Gardner, H. P. Graf, D. Henderson, R. E. Howard, W. Hubbard,
L. D. Jackel, H. S. Baird, and I. Guyon** (Advances in NIPS 1, 1988) — **Yann
LeCun is not among the authors**. LeCun is first author only on the *1989* Neural
Computation paper "Backpropagation Applied to Handwritten Zip Code Recognition"
([[claim-lecun-1989-first-practical-recognition]]). Several names overlap (Denker,
Henderson, Howard, Hubbard, Jackel) — the same group — which is how retellings
slide "Denker et al. 1988" into "LeCun 1988."

**Status: contested — the authorship half is wrong, the mechanism half is right.**
Read directly, the 1988 Denker paper does use hand-designed feature detectors:
its 49 "feature extractor templates" (7×7 pixels) are explicitly designed by the
authors — "Feature b is designed to detect the right-hand end of (approximately)
horizontal strokes… the image must be able to touch the 'should be ON' pixels…
without touching the surrounding horseshoe-shaped collection of 'must be OFF'
pixels" — a hand-crafted detector, not a learned kernel. So the field lore is
right that the 1988 system's feature stage was hand-designed and the 1989 system
learned its kernels by [[entity-backpropagation|backpropagation]]; it is wrong only in attaching LeCun's name
to the 1988 paper.

This is the same priority-slide the vault tracks for
[[claim-fukushima-1979-neocognitron-first-cnn]] — a group result migrating to the
one remembered name. See [[moc-backpropagation-origins]].

**A wider pattern this transition instantiates.** The 1988 hand-designed →
1989 backprop-learned kernel shift is a two-year-early case of what Rich
Sutton's 2019 essay "The Bitter Lesson" would later name as a 70-year AI
pattern — hand-engineered human knowledge wins short-term, then loses to
general methods that leverage computation — and Sutton cites vision
(hand-designed edges/SIFT features vs. learned convolution) as one of his
four examples. That essay is also the hinge to
[[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]]:
Sutton says such methods "scale arbitrarily," but neural-scaling-law exponents
(Kaplan et al. 2020, α≈0.05–0.095) show the scaling that wins is real but
logarithmically diminishing — the same sub-linear brake that note documents
for evolution and idea-production. See 2026-07-27-hop-bitter-lesson-scaling-brake
for the full chain — promoted 2026-07-28 as
[[claim-sutton-2019-bitter-lesson-names-pattern-silent-on-rate]] and
[[claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns]].

> [!note] Seek's commentary:
> The failure mode here is different from the Werbos case even though both end in "one name for many people." There, a citation vacuum ([[claim-rhw-1986-reference-list-four-works]]) pulled a single inventor in from nowhere; here the *fact* is right (1988 hand-designed, 1989 learned) and only the name is wrong — Denker et al.'s paper drifts onto LeCun because fame works as a citation attractor, the same slide that moved the neocognitron toward its one remembered author ([[claim-fukushima-1979-neocognitron-first-cnn]]). Vacuum-filling and name-magnetism are two distinct engines producing the same distortion; worth keeping separate, because you'd correct them differently.
> — Seek
