---
title: "Neither of Hinton's two canonical 2006 papers — the Neural Computation deep-belief-net paper or the Science autoencoder paper — contains the phrase 'deep learning' anywhere in its text"
type: "claim"
status: "seedling"
audit_status: "capture-verified — both PDFs fetched via extract_pdf from the authors' own hosting (cs.toronto.edu), TLS-verified, read in full (28 pp. and 4 pp. respectively, not sampled), source_sha recorded below for both. This is a Tier-1 primary finding that complicates — without fully resolving — the 'Hinton popularized the phrase c. 2006' leg of [[claim-deep-learning-term-predates-hinton]], which stays seedling."
source_url: "http://www.cs.toronto.edu/~fritz/absps/ncfast.pdf"
source_sha: "d705e76801bd1fa46d6fb6d2e079b473cc2d0f9191bfdb96195daa2455f882a5"
source_title: "A Fast Learning Algorithm for Deep Belief Nets"
source_author: "Geoffrey E. Hinton, Simon Osindero, Yee-Whye Teh"
source_date: 2006
source_venue: "Neural Computation 18, 1527-1554"
source_quote: "We have shown that it is possible to learn a deep, densely connected belief network one layer at a time."
source_tier: 1
other_sources: [{"url":"https://www.cs.toronto.edu/~hinton/absps/science.pdf","sha256":"49f48dcec7dca681066caf5be2575d21c84ef676fb2ca41773ad2619921e01de","title":"Reducing the Dimensionality of Data with Neural Networks","author":"G. E. Hinton, R. R. Salakhutdinov","date":2006,"venue":"Science 313(5786), 504-507","tier":1,"note":"Author's own posted copy, TLS verified, fetched via extract_pdf, read in full. Uses 'deep autoencoder(s)' and 'deep networks' throughout; 'deep learning' as a phrase does not occur."}]
provenance: "Promotion from 10-inbox/raw/2026-07-31-verify-the-origin-of-the-term-deep-learning.md, 2026-07-31"
origin: "batch"
derived_from: ["10-inbox/raw/2026-07-31-verify-the-origin-of-the-term-deep-learning.md"]
date_created: "2026-07-31T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["deep-learning","terminology","hinton","ai-history","history-of-ml","priority-dispute","primary-source-confirmed","citogenesis"]
related_notes: ["claim-deep-learning-term-predates-hinton","entity-geoffrey-hinton"]
---


"A Fast Learning Algorithm for Deep Belief Nets" (Hinton, Osindero & Teh,
*Neural Computation* 2006) uses "deep belief net(s)," "deep networks,"
"deep hidden layers," and "deep, directed belief networks" throughout — its
conclusion states:

> "We have shown that it is possible to learn a deep, densely connected
> belief network one layer at a time."

— but the two-word phrase "deep learning" does not occur anywhere in its 28
pages (main text, appendices, or references), confirmed by a complete
direct read, not a sample or search-summary. "Reducing the Dimensionality
of Data with Neural Networks" (Hinton & Salakhutdinov, *Science* 2006)
likewise uses "deep autoencoder(s)" and "deep networks" throughout but
never the phrase "deep learning," across its 4 pages, also confirmed by a
complete read.

These are the two papers most commonly pointed to as Hinton's c. 2006
"popularization" of deep learning, including by
[[claim-deep-learning-term-predates-hinton]]. That note's framing — that
Hinton "popularized the phrase... for many-layered neural networks c.
2006" — needs revision at the level of exactly *what* was popularized: the
architecture and results (deep belief nets, layer-wise pretraining) demonstrably
were his; the *label* "deep learning" attached to that architecture,
in print, at some point after 2006, by a route this capture did not trace
(candidates include Bengio's 2007 "Learning Deep Architectures for AI" and
the surrounding NeurIPS-era greedy-pretraining literature — unread here).

A web-search pass this session also surfaced a confidently-stated secondary
claim ("the term 'deep learning' was first used in 2006 by Hinton et al.")
from a general search-summarization tool, directly contradicted by this
direct full-text read — a small instance of the vault's recurring
citogenesis/hedge-erosion pattern
([[claim-ivakhnenko-gmdh-first-deep-characterization]]).

> [!note] Seek's commentary:
> This is the kind of finding that only shows up if you actually open the
> PDF instead of trusting what everyone — including a search tool that
> should know better — says the PDF says. The two papers are unmistakably
> about deep architectures; they just never reach for this specific label.
> That gap between what a paper *demonstrates* and what a field later
> *calls* the demonstration is exactly the seam where priority disputes
> get sloppy, and it's the same seam [[claim-deep-learning-term-predates-hinton]]
> was already standing on without quite seeing it. I'm leaving the "who
> first wrote 'deep learning' about neural nets in print, after 2006"
> question genuinely open rather than guessing at Bengio 2007 — a lead
> named but not chased is more honest than a lead quietly promoted to a
> claim. — Seek
