Verify the origin of the term 'deep learning' against primaries: Dechter (1986), Aizenberg et al. (2000), and the Hinton-era popularization (c. 2006)
This capture follows the routed question
question-verify-deep-learning-term-origin-dechter-aizenberg (itself raised
against claim-deep-learning-term-predates-hinton, which currently rests on
Tier-4 Wikipedia). Three primaries were targeted: Dechter (1986), Aizenberg et
al. (2000), and Hinton's c. 2006 papers. Two of the three were directly read in
full via extract_pdf; the third (Aizenberg) is paywalled and was not
independently confirmed.
Bottom line on the central question: the Dechter (1986) leg of the Wikipedia chain is now CONFIRMED at Tier 1 — she genuinely uses "deep learning" and "shallow learning" as terms of art, though for search depth in constraint-satisfaction backtracking, not neural networks. The Hinton-era leg is COMPLICATED, not confirmed as stated: neither of the two papers most commonly cited as the c. 2006 "popularization" actually contains the phrase "deep learning" anywhere in their text. The Aizenberg leg remains UNVERIFIED — the primary book could not be read, and the claim traces to a single page-number-free Wikipedia citation.
Claim: Rina Dechter's 1986 AAAI paper genuinely uses the phrase "deep learning" (paired with "shallow learning") as a technical term — but for controlling search depth in constraint-satisfaction backtracking, not for neural networks
Claim type: historical/terminological, escalated per the sourcing floor
because it is the load-bearing point of a priority dispute. Tier 1-2
required — achieved Tier 1 (direct full-text read of the primary via
extract_pdf, own AAAI-proceedings venue).
Dechter, "Learning While Searching in Constraint-Satisfaction-Problems," AAAI-86 (pp. 178-183), studies backtracking search enhanced by recording constraints learned at "dead-ends." She defines "deep" vs. "shallow" learning as a parameter of how much conflict-set information to record, and reports:
"Discovering all minimal conflict-sets amounts to acquiring all the possible information out of a dead-end. Yet, such deep learning may require considerable amount of work."
and, in the experimental results:
"in most cases both performance measures improve as we move from shallow learning to deep learning and from first-order to second-order."
Table columns in the paper are literally labeled "DF = Deep-First-Order" and "DS = Deep-Second-order," contrasted with "SF = Shallow-First-Order" and "SS = Shallow-Second-Order." The conclusion names "deep-second-order learning" as "the strongest form of learning we have tested." This is Wikipedia's claim confirmed at the primary level: the phrase is genuinely hers, in 1986, and it is genuinely about a notion of depth of search/recording, not about many-layered neural architectures — the referent Wikipedia's own gloss ("in the context of Boolean threshold neurons" is Aizenberg's clause, not Dechter's) already implies but does not spell out.
Provenance: source_url https://cdn.aaai.org/AAAI/1986/AAAI86-029.pdf,
sha256 59f6c1ca738a9acb5bc49985da12fae76ce2226e65899e18cbb281669f71c4f0
(TLS verified), fetched via extract_pdf, full 6-page text read directly.
Claim: Neither of the two papers most commonly cited as Hinton's c. 2006 "popularization" of deep learning — Hinton, Osindero & Teh (Neural Computation, 2006) and Hinton & Salakhutdinov (Science, 2006) — contains the phrase "deep learning" anywhere in its text
Claim type: historical/bibliographic (a full-text-presence fact, directly
checkable), escalated per the sourcing floor because it is surprising and
load-bearing for a priority dispute. Tier 1-2 required — achieved Tier 1
(both PDFs fetched via extract_pdf from the authors' own hosting at
cs.toronto.edu; both read in full, page by page, not sampled).
"A Fast Learning Algorithm for Deep Belief Nets" (Hinton, Osindero & Teh 2006) uses "deep belief net(s)," "deep networks," "deep hidden layers," and "deep, directed belief networks" throughout — e.g., its conclusion:
"We have shown that it is possible to learn a deep, densely connected belief network one layer at a time."
but the two-word phrase "deep learning" does not occur anywhere in its 28 pages (main text, appendices, or references), confirmed by a complete read.
"Reducing the Dimensionality of Data with Neural Networks" (Hinton & Salakhutdinov, Science 2006) likewise uses "deep autoencoder(s)" and "deep networks" throughout — e.g.:
"It has been obvious since the 1980s that backpropagation through deep autoencoders would be very effective for nonlinear dimensionality reduction, provided that computers were fast enough, data sets were big enough, and the initial weights were close enough to a good solution."
but again, "deep learning" as a phrase does not appear anywhere in its 4 pages, confirmed by a complete read.
This directly complicates claim-deep-learning-term-predates-hinton's framing that Hinton "popularized the phrase" for many-layered neural networks c. 2006: the two papers that note cites (and that the field usually points to as the 2006 moment) demonstrate the architecture — "deep belief nets," "deep autoencoders" — but not the label "deep learning" itself. A web-search pass this session surfaced a secondary-source claim ("the term 'deep learning' was first used in 2006 by Hinton et al.") that a general search-summarization tool asserted confidently — directly contradicted by the primary full-text read here. That contradiction is itself worth flagging as an instance of the vault's recurring citogenesis/hedge-erosion pattern (claim-ivakhnenko-gmdh-first-deep-characterization, myth-amari-first-sgd-mlp): a plausible-sounding secondary claim that does not survive a direct read of the text it's supposedly about. Precisely when "deep learning" (the label) came to be attached to Hinton-style deep networks in the community's usage is not established by this capture and is left as a further lead below.
Provenance:
- Hinton, Osindero & Teh 2006 — source_url
http://www.cs.toronto.edu/~fritz/absps/ncfast.pdf, sha256d705e76801bd1fa46d6fb6d2e079b473cc2d0f9191bfdb96195daa2455f882a5, TLS verified, fetched viaextract_pdf, full text read. - Hinton & Salakhutdinov 2006 — source_url
https://www.cs.toronto.edu/~hinton/absps/science.pdf, sha25649f48dcec7dca681066caf5be2575d21c84ef676fb2ca41773ad2619921e01de, TLS verified, fetched viaextract_pdf, full text read.
Claim: The Aizenberg, Aizenberg & Vandewalle (2000) attribution — that "deep learning" was applied to artificial neural networks in the context of Boolean threshold neurons — rests entirely on an unverified Wikipedia citation and could not be confirmed against the primary text
Claim type: historical/terminological, load-bearing for the priority
dispute, therefore escalated. Tier 1-2 required — NOT achieved.
[unverified-historical — needs primary].
Wikipedia's "Deep learning" article cites Multi-Valued and Universal Binary
Neurons: Theory, Learning and Applications (Aizenberg, Aizenberg &
Vandewalle, Kluwer, 2000, DOI 10.1007/978-1-4757-3115-6) for this claim, but
the wikitext citation carries no page number — confirmed by fetching the raw
wikitext directly (action=raw) and reading the {{cite book}} template for
ref MV_1. Attempts to reach the primary text this session:
- Springer's book page (
link.springer.com/book/10.1007/978-1-4757-3115-6) redirected to an authentication/paywall gate (idp.springer.com); no full-text or preview access obtained. - A Google Books search-within-book query for "deep learning" against the book's Google Books ID returned no visible snippet or page number.
- General web search repeatedly surfaced the same sentence structure ("the term deep learning was first introduced to the realm of artificial neural networks by Igor Aizenberg... in the context of Boolean threshold neurons") verbatim or near-verbatim across multiple secondary pages — consistent with all of them deriving from the same Wikipedia sentence rather than independently checking the book, the same citogenesis risk the vault has already flagged for the Amari/Wikipedia and Ivakhnenko/Schmidhuber threads (claim-ivakhnenko-gmdh-first-deep-characterization).
The claim is recorded here as a lead, not as a verified fact: it may well be true (the book is a real, peer-reviewed-adjacent Kluwer monograph by named authors, and Aizenberg's group did work on multi-valued/Boolean-threshold neurons in this period), but no independent primary confirmation was obtained after genuine effort this session.
Provenance: source is Wikipedia's own citation record (Tier 4) for the underlying book (Tier unconfirmed — could not be read).
Central question status
Does the term "deep learning" originate with Dechter (1986), get applied to neural networks by Aizenberg et al. (2000), and get popularized by Hinton c. 2006?
Partially confirmed, partially complicated, partially unverified — not a
clean yes. Dechter's 1986 usage is now Tier-1-confirmed as genuine, in the
constraint-satisfaction-search sense. The Hinton c. 2006 "popularization" is
complicated by a primary-text finding this capture surfaces: the two papers
usually cited for that moment do not use the phrase "deep learning" at all,
only "deep belief nets"/"deep autoencoders"/"deep networks" — so the
coinage-vs-popularization frame in claim-deep-learning-term-predates-hinton
needs revision at the level of exactly what was popularized and when the
label itself (as opposed to the architecture) caught on. The Aizenberg leg
remains [unverified-historical — needs primary] — the book was not
independently readable this session.
Further leads
- When did "deep learning" as a literal two-word label first get attached to Hinton/Bengio/LeCun-style many-layered neural nets in print? Bengio's 2007 "Learning Deep Architectures for AI" technical report and the Bengio/Ranzato/Larochelle 2007 NeurIPS-era papers on greedy layer-wise pretraining are candidates worth a direct full-text check, not yet done here.
- A library/interlibrary-loan or borrow-access read of Aizenberg, Aizenberg & Vandewalle (2000) itself (e.g. via Internet Archive borrow, if available) would close the unverified leg.
- Rina Dechter's own later work and CV (UC Irvine) might state, in her own words, whether she is aware of / claims priority for the "deep learning" phrase — an entity-level check, not done here.
- The general-purpose web-search tool's confident-but-wrong claim ("the term was first used in 2006 by Hinton et al.") is itself a small case study in search-summary unreliability, worth a note to future researchers using the same tool for historical-priority questions.
Entity candidates
- Rina Dechter — person — the earliest confirmed primary user of the literal phrase "deep learning" (1986); no entity page exists yet for her in the vault despite being the foundational figure this whole priority dispute measures Hinton and Aizenberg against. Flagged first per the vault's known blind spot for under-flagging the older figure a priority claim rests on.
- Igor Aizenberg — person — credited (via unverified secondary attribution) with applying "deep learning" to artificial neural networks in 2000; no entity page exists yet; the claim resting on him is currently unverifiable against the primary.
- Naum N. Aizenberg — person — co-author of the 2000 book (Igor's father, per publicly known biographical context; not verified in this session) — flagged for completeness alongside Igor.
- Joos P. L. Vandewalle — person — third co-author of the 2000 book.
- Geoffrey Hinton — person — already has vault presence (entity-geoffrey-hinton referenced elsewhere); central to the "popularizer" side of this dispute; this capture's finding (his 2006 papers don't use the phrase) bears directly on how his entity note should characterize his role.
- Simon Osindero — person — co-author, Hinton, Osindero & Teh 2006.
- Yee-Whye Teh — person — co-author, Hinton, Osindero & Teh 2006.
- Ruslan Salakhutdinov — person — co-author, Hinton & Salakhutdinov 2006 Science paper.
- deep belief net — concept/term — the actual phrase used in the 2006 Hinton-era papers, distinct from "deep learning"; may warrant its own concept note given this capture's finding that the two terms are not interchangeable in the primary texts.
- citogenesis (Wikipedia-claim self-reinforcement) — concept — the pattern this capture repeatedly ran into (Aizenberg attribution, the "Hinton first used it in 2006" search-tool claim); already tracked elsewhere in the vault via the Amari/Ivakhnenko threads and may deserve its own consolidated concept note.
Source
claude-sonnet-5 · Batch run 2026-07-31, routed from 50-questions/question-verify-deep-learning-term-origin-dechter-aizenberg.md. Researched via direct extract_pdf reads of the Dechter (1986) AAAI proceedings PDF and both canonical Hinton 2006 papers (Neural Computation and Science); web search for the Aizenberg et al. (2000) book, which is paywalled (Springer/Kluwer) and could not be read in full text. · raw markdown