---
id: "20260731-0201-verify-the-origin-of"
title: "Verify the origin of the term 'deep learning' against primaries: Dechter (1986), Aizenberg et al. (2000), and the Hinton-era popularization (c. 2006)"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-dechter-1986-deep-learning-shallow-learning-csp-search-depth.md","30-notes/claim-hinton-2006-papers-omit-deep-learning-phrase.md","40-entities/entity-rina-dechter.md","40-entities/entity-citogenesis.md"]
updated_existing: ["40-entities/entity-geoffrey-hinton.md — dated append line (2026-07-31) noting his 2006 papers omit the phrase 'deep learning'","50-questions/question-verify-deep-learning-term-origin-dechter-aizenberg.md — progress line added, ruled partially answered, left status: open (Aizenberg leg still unresolved)"]
not_promoted: ["The Aizenberg, Aizenberg & Vandewalle (2000) attribution claim — not promoted as a standalone claim-note. It remains [unverified-historical — needs primary] (paywalled, no page number in the source citation); that non-finding is already carried by claim-deep-learning-term-predates-hinton.md's existing audit_status flag and by the progress-line update to the routed question. A third note restating 'still unverified' would not add permanent content beyond what those two already record.","Further leads (Bengio 2007 'Learning Deep Architectures for AI' read; Dechter's own CV/priority stance; a meta-note on the search-summarization tool's fabricated-sounding 'Hinton coined it in 2006' claim) — left as leads inside the routed question and the new claim-notes' commentary, not opened as new 50-questions/ entries. Per question-intake discipline, none of the kept claims rest on these; they are nice-to-verify next steps, not load-bearing doubts, so no new promise was made to the question pile.","Entity candidates Igor Aizenberg, Naum N. Aizenberg, Joos P. L. Vandewalle, Simon Osindero, Yee-Whye Teh, Ruslan Salakhutdinov — not promoted to entity hubs. Igor Aizenberg's vault relevance is entirely contingent on a claim that could not be verified this session (UNSURE per the entity-promotion test, bias against the flood); the others are one-off co-author mentions, not recurring figures.","'Deep belief net' as a concept candidate — not promoted to its own entity hub. Real and established in the field, but this session is its first appearance anywhere in the vault (not yet recurring or load-bearing within the vault itself), so it stays a mention inside claim-hinton-2006-papers-omit-deep-learning-phrase.md rather than a stub page."]
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-07-31T00:00:00.000Z"
provenance: "Batch run 2026-07-31, routed from 50-questions/question-verify-deep-learning-term-origin-dechter-aizenberg.md. Researched via direct extract_pdf reads of the Dechter (1986) AAAI proceedings PDF and both canonical Hinton 2006 papers (Neural Computation and Science); web search for the Aizenberg et al. (2000) book, which is paywalled (Springer/Kluwer) and could not be read in full text."
derived_from: []
tags: ["deep-learning","terminology","priority-dispute","ai-history","history-of-ml","dechter","aizenberg","hinton","primary-source-confirmed"]
source_url: "https://cdn.aaai.org/AAAI/1986/AAAI86-029.pdf"
source_sha: "59f6c1ca738a9acb5bc49985da12fae76ce2226e65899e18cbb281669f71c4f0"
source_title: "Learning While Searching in Constraint-Satisfaction-Problems"
source_author: "Rina Dechter"
source_date: 1986
source_venue: "Proceedings of the Fifth AAAI National Conference on Artificial Intelligence (AAAI-86), pp. 178-183"
source_tier: 1
other_sources: [{"url":"http://www.cs.toronto.edu/~fritz/absps/ncfast.pdf","sha256":"d705e76801bd1fa46d6fb6d2e079b473cc2d0f9191bfdb96195daa2455f882a5","title":"A Fast Learning Algorithm for Deep Belief Nets","author":"Geoffrey E. Hinton, Simon Osindero, Yee-Whye Teh","date":2006,"venue":"Neural Computation 18, 1527-1554","tier":1,"note":"Author's own posted copy (cs.toronto.edu), read in full (28 pp.). TLS verified."},{"url":"https://www.cs.toronto.edu/~hinton/absps/science.pdf","sha256":"49f48dcec7dca681066caf5be2575d21c84ef676fb2ca41773ad2619921e01de","title":"Reducing the Dimensionality of Data with Neural Networks","author":"G. E. Hinton, R. R. Salakhutdinov","date":2006,"venue":"Science 313(5786), 504-507","tier":1,"note":"Author's own posted copy (cs.toronto.edu), read in full. TLS verified."},{"url":"https://en.wikipedia.org/wiki/Deep_learning","title":"Deep learning","date":"accessed 2026-07-31","venue":"Wikipedia","tier":4,"note":"Source of the Dechter/Aizenberg/Hinton attribution chain this capture set out to verify. Its own wikitext citations for the Aizenberg claim point only to the book (DOI 10.1007/978-1-4757-3115-6), no page number given."},{"url":"https://link.springer.com/book/10.1007/978-1-4757-3115-6","title":"Multi-Valued and Universal Binary Neurons: Theory, Learning and Applications","author":"Igor Aizenberg, Naum N. Aizenberg, Joos P. L. Vandewalle","date":2000,"venue":"Kluwer Academic Publishers (Springer)","tier":"unconfirmed — paywalled, full text not obtained","note":"Springer redirected to an authentication/paywall page (idp.springer.com); no accessible preview via Google Books search turned up a 'deep learning' snippet with page number."}]
---


This capture follows the routed question
[[question-verify-deep-learning-term-origin-dechter-aizenberg]] (itself raised
against [[claim-deep-learning-term-predates-hinton]], which currently rests on
Tier-4 Wikipedia). Three primaries were targeted: Dechter (1986), Aizenberg et
al. (2000), and Hinton's c. 2006 papers. Two of the three were directly read in
full via `extract_pdf`; the third (Aizenberg) is paywalled and was not
independently confirmed.

**Bottom line on the central question:** the Dechter (1986) leg of the
Wikipedia chain is now CONFIRMED at Tier 1 — she genuinely uses "deep learning"
and "shallow learning" as terms of art, though for search depth in
constraint-satisfaction backtracking, not neural networks. The Hinton-era leg
is COMPLICATED, not confirmed as stated: neither of the two papers most
commonly cited as the c. 2006 "popularization" actually contains the phrase
"deep learning" anywhere in their text. The Aizenberg leg remains UNVERIFIED —
the primary book could not be read, and the claim traces to a single
page-number-free Wikipedia citation.

---

## Claim: Rina Dechter's 1986 AAAI paper genuinely uses the phrase "deep learning" (paired with "shallow learning") as a technical term — but for controlling search depth in constraint-satisfaction backtracking, not for neural networks

**Claim type:** historical/terminological, escalated per the sourcing floor
because it is the load-bearing point of a priority dispute. **Tier 1-2
required — achieved Tier 1** (direct full-text read of the primary via
`extract_pdf`, own AAAI-proceedings venue).

Dechter, "Learning While Searching in Constraint-Satisfaction-Problems," AAAI-86
(pp. 178-183), studies backtracking search enhanced by recording constraints
learned at "dead-ends." She defines "deep" vs. "shallow" learning as a
parameter of how much conflict-set information to record, and reports:

> "Discovering all minimal conflict-sets amounts to acquiring all the possible
> information out of a dead-end. Yet, such deep learning may require
> considerable amount of work."

and, in the experimental results:

> "in most cases both performance measures improve as we move from shallow
> learning to deep learning and from first-order to second-order."

Table columns in the paper are literally labeled "DF = Deep-First-Order" and
"DS = Deep-Second-order," contrasted with "SF = Shallow-First-Order" and "SS =
Shallow-Second-Order." The conclusion names "deep-second-order learning" as
"the strongest form of learning we have tested." This is Wikipedia's claim
confirmed at the primary level: the phrase is genuinely hers, in 1986, and it
is genuinely about a notion of *depth of search/recording*, not about
many-layered neural architectures — the referent Wikipedia's own gloss
("in the context of Boolean threshold neurons" is Aizenberg's clause, not
Dechter's) already implies but does not spell out.

**Provenance:** source_url `https://cdn.aaai.org/AAAI/1986/AAAI86-029.pdf`,
sha256 `59f6c1ca738a9acb5bc49985da12fae76ce2226e65899e18cbb281669f71c4f0`
(TLS verified), fetched via `extract_pdf`, full 6-page text read directly.

---

## Claim: Neither of the two papers most commonly cited as Hinton's c. 2006 "popularization" of deep learning — Hinton, Osindero & Teh (*Neural Computation*, 2006) and Hinton & Salakhutdinov (*Science*, 2006) — contains the phrase "deep learning" anywhere in its text

**Claim type:** historical/bibliographic (a full-text-presence fact, directly
checkable), escalated per the sourcing floor because it is surprising and
load-bearing for a priority dispute. **Tier 1-2 required — achieved Tier 1**
(both PDFs fetched via `extract_pdf` from the authors' own hosting at
cs.toronto.edu; both read in full, page by page, not sampled).

"A Fast Learning Algorithm for Deep Belief Nets" (Hinton, Osindero & Teh 2006)
uses "deep belief net(s)," "deep networks," "deep hidden layers," and "deep,
directed belief networks" throughout — e.g., its conclusion:

> "We have shown that it is possible to learn a deep, densely connected belief
> network one layer at a time."

but the two-word phrase "deep learning" does not occur anywhere in its 28
pages (main text, appendices, or references), confirmed by a complete read.

"Reducing the Dimensionality of Data with Neural Networks" (Hinton &
Salakhutdinov, *Science* 2006) likewise uses "deep autoencoder(s)" and "deep
networks" throughout — e.g.:

> "It has been obvious since the 1980s that backpropagation through deep
> autoencoders would be very effective for nonlinear dimensionality
> reduction, provided that computers were fast enough, data sets were big
> enough, and the initial weights were close enough to a good solution."

but again, "deep learning" as a phrase does not appear anywhere in its 4
pages, confirmed by a complete read.

This directly complicates [[claim-deep-learning-term-predates-hinton]]'s
framing that Hinton "popularized the phrase" for many-layered neural networks
c. 2006: the two papers that note cites (and that the field usually points to
as the 2006 moment) demonstrate the *architecture* — "deep belief nets,"
"deep autoencoders" — but not the *label* "deep learning" itself. A web-search
pass this session surfaced a secondary-source claim ("the term 'deep learning'
was first used in 2006 by Hinton et al.") that a general search-summarization
tool asserted confidently — directly contradicted by the primary full-text
read here. That contradiction is itself worth flagging as an instance of the
vault's recurring citogenesis/hedge-erosion pattern
([[claim-ivakhnenko-gmdh-first-deep-characterization]],
[[myth-amari-first-sgd-mlp]]): a plausible-sounding secondary claim that does
not survive a direct read of the text it's supposedly about. Precisely when
"deep learning" (the label) came to be attached to Hinton-style deep networks
in the community's usage is not established by this capture and is left as a
further lead below.

**Provenance:**
- Hinton, Osindero & Teh 2006 — source_url `http://www.cs.toronto.edu/~fritz/absps/ncfast.pdf`, sha256 `d705e76801bd1fa46d6fb6d2e079b473cc2d0f9191bfdb96195daa2455f882a5`, TLS verified, fetched via `extract_pdf`, full text read.
- Hinton & Salakhutdinov 2006 — source_url `https://www.cs.toronto.edu/~hinton/absps/science.pdf`, sha256 `49f48dcec7dca681066caf5be2575d21c84ef676fb2ca41773ad2619921e01de`, TLS verified, fetched via `extract_pdf`, full text read.

---

## Claim: The Aizenberg, Aizenberg & Vandewalle (2000) attribution — that "deep learning" was applied to artificial neural networks in the context of Boolean threshold neurons — rests entirely on an unverified Wikipedia citation and could not be confirmed against the primary text

**Claim type:** historical/terminological, load-bearing for the priority
dispute, therefore escalated. **Tier 1-2 required — NOT achieved.**
`[unverified-historical — needs primary]`.

Wikipedia's "Deep learning" article cites *Multi-Valued and Universal Binary
Neurons: Theory, Learning and Applications* (Aizenberg, Aizenberg &
Vandewalle, Kluwer, 2000, DOI 10.1007/978-1-4757-3115-6) for this claim, but
the wikitext citation carries no page number — confirmed by fetching the raw
wikitext directly (`action=raw`) and reading the `{{cite book}}` template for
ref `MV_1`. Attempts to reach the primary text this session:

- Springer's book page (`link.springer.com/book/10.1007/978-1-4757-3115-6`)
  redirected to an authentication/paywall gate (`idp.springer.com`); no
  full-text or preview access obtained.
- A Google Books search-within-book query for "deep learning" against the
  book's Google Books ID returned no visible snippet or page number.
- General web search repeatedly surfaced the *same* sentence structure
  ("the term deep learning was first introduced to the realm of artificial
  neural networks by Igor Aizenberg... in the context of Boolean threshold
  neurons") verbatim or near-verbatim across multiple secondary
  pages — consistent with all of them deriving from the same Wikipedia
  sentence rather than independently checking the book, the same
  citogenesis risk the vault has already flagged for the Amari/Wikipedia and
  Ivakhnenko/Schmidhuber threads
  ([[claim-ivakhnenko-gmdh-first-deep-characterization]]).

The claim is recorded here as a lead, not as a verified fact: it may well be
true (the book is a real, peer-reviewed-adjacent Kluwer monograph by named
authors, and Aizenberg's group did work on multi-valued/Boolean-threshold
neurons in this period), but no independent primary confirmation was obtained
after genuine effort this session.

**Provenance:** source is Wikipedia's own citation record (Tier 4) for the
underlying book (Tier unconfirmed — could not be read).

---

## Central question status

**Does the term "deep learning" originate with Dechter (1986), get applied to
neural networks by Aizenberg et al. (2000), and get popularized by Hinton c.
2006?**

**Partially confirmed, partially complicated, partially unverified — not a
clean yes.** Dechter's 1986 usage is now Tier-1-confirmed as genuine, in the
constraint-satisfaction-search sense. The Hinton c. 2006 "popularization" is
complicated by a primary-text finding this capture surfaces: the two papers
usually cited for that moment do not use the phrase "deep learning" at all,
only "deep belief nets"/"deep autoencoders"/"deep networks" — so the
coinage-vs-popularization frame in [[claim-deep-learning-term-predates-hinton]]
needs revision at the level of exactly *what* was popularized and *when* the
label itself (as opposed to the architecture) caught on. The Aizenberg leg
remains `[unverified-historical — needs primary]` — the book was not
independently readable this session.

---

## Further leads

- When did "deep learning" as a literal two-word label first get attached to Hinton/Bengio/LeCun-style many-layered neural nets in print? Bengio's 2007 "Learning Deep Architectures for AI" technical report and the Bengio/Ranzato/Larochelle 2007 NeurIPS-era papers on greedy layer-wise pretraining are candidates worth a direct full-text check, not yet done here.
- A library/interlibrary-loan or borrow-access read of Aizenberg, Aizenberg & Vandewalle (2000) itself (e.g. via Internet Archive borrow, if available) would close the unverified leg.
- Rina Dechter's own later work and CV (UC Irvine) might state, in her own words, whether she is aware of / claims priority for the "deep learning" phrase — an entity-level check, not done here.
- The general-purpose web-search tool's confident-but-wrong claim ("the term was first used in 2006 by Hinton et al.") is itself a small case study in search-summary unreliability, worth a note to future researchers using the same tool for historical-priority questions.

---

## Entity candidates

- Rina Dechter — person — the earliest confirmed primary user of the literal phrase "deep learning" (1986); no entity page exists yet for her in the vault despite being the foundational figure this whole priority dispute measures Hinton and Aizenberg against. Flagged first per the vault's known blind spot for under-flagging the older figure a priority claim rests on.
- Igor Aizenberg — person — credited (via unverified secondary attribution) with applying "deep learning" to artificial neural networks in 2000; no entity page exists yet; the claim resting on him is currently unverifiable against the primary.
- Naum N. Aizenberg — person — co-author of the 2000 book (Igor's father, per publicly known biographical context; not verified in this session) — flagged for completeness alongside Igor.
- Joos P. L. Vandewalle — person — third co-author of the 2000 book.
- Geoffrey Hinton — person — already has vault presence ([[entity-geoffrey-hinton]] referenced elsewhere); central to the "popularizer" side of this dispute; this capture's finding (his 2006 papers don't use the phrase) bears directly on how his entity note should characterize his role.
- Simon Osindero — person — co-author, Hinton, Osindero & Teh 2006.
- Yee-Whye Teh — person — co-author, Hinton, Osindero & Teh 2006.
- Ruslan Salakhutdinov — person — co-author, Hinton & Salakhutdinov 2006 Science paper.
- deep belief net — concept/term — the actual phrase used in the 2006 Hinton-era papers, distinct from "deep learning"; may warrant its own concept note given this capture's finding that the two terms are not interchangeable in the primary texts.
- citogenesis (Wikipedia-claim self-reinforcement) — concept — the pattern this capture repeatedly ran into (Aizenberg attribution, the "Hinton first used it in 2006" search-tool claim); already tracked elsewhere in the vault via the Amari/Ivakhnenko threads and may deserve its own consolidated concept note.
