Widrow's group spent the 1960s unable to train a genuinely multilayer network and left neural-net research until backpropagation reached Widrow in the mid-1980s
After Widrow and Hoff trained the single-layer Adaline by gradient descent (LMS) in the early 1960s, the group spent years trying to extend learning to a second, hidden layer and could not. The reason is stated plainly in Widrow's own retrospective: the original Madalines used hard-limiting quantizers (signums), whose lack of a usable derivative blocks the backward error signal that a differentiable "sigmoid" network permits (claim-madaline-rule-i-first-layer-trainable-second-fixed). That Widrow's group was still, in 1987, "looking back at the original Madaline I algorithm with the goal of developing a new technique that could adapt multiple layers" is itself evidence that the 1960s effort never produced one.
Unable to train hidden layers, Widrow left neural-network research for adaptive signal processing (adaptive antennas, noise cancelling, seismic processing); Hoff went to Intel. Per Yuxi Liu's footnoted history and Wikipedia's Widrow biography, he returned only after encountering backpropagation around a 1985 Snowbird, Utah conference — the specific date and venue rest on secondary sources, not on Widrow's own oral history, now that it has been read directly (see Correction history below), and are flagged accordingly above.
A second 1960s multilayer attempt, independent of Madaline Rule I, is now primary-documented: Widrow's 1966 "Bootstrap Learning" paper describes a preliminary, non-gradient, whole-network reinforcement scheme for multilayer training of adaptive threshold elements (claim-widrow-1966-bootstrap-learning-preliminary-multilayer-attempt) — a second, earlier witness to the same "unsuccessful attempts" fact this note already carries from the 1990 retrospective.
This is the empirical companion to the Perceptrons stall narrative (claim-perceptrons-multilayer-sterile-was-conjecture, myth-perceptrons-book-killed-connectionism): Minsky and Papert conjectured multilayer training was a dead end; Widrow's group lived the dead end, for the concrete reason that gradient descent needs a differentiable nonlinearity the Madaline hardware did not have. The barrier fell only when the differentiable-sigmoid route was demonstrated to learn useful hidden representations (claim-rhw-1986-demonstration-not-invention). See backpropagation-gap and moc-backpropagation-origins.
A gradient-free counterpoint: persistent homology characterizes a network's structure with no derivative at all — the exact constraint that defeated Widrow's hard-limiters — and now bounds neural-network generalization from the training trajectory. See observation-persistent-homology-gradient-free-bridge-widrow-byzantine.
Correction history.
- 2026-09-05 — question-verify-widrow-1965-66-papers-multilayer-attempt's oral-history verification landed: the full 1997 ETHW transcript (Widrow, interviewed by Andrew Goldstein, 14 March 1997) was read directly, end to end, this session (
source_sha: 035ee2b8effe2ee6903c9ebd5e08742824cc6ca596a02b93d3f034fa18de13b4). It does not mention 1985, Snowbird, backpropagation, sigmoids, or a return to neural-network research at all — the "1985 Snowbird" story does not trace to this document, contrary to this note's earlier framing of it as merely "unread." The closely related "sharp quantizers / need sigmoids / no one knew anything about it" quote is real but traces instead to Talking Nets: An Oral History of Neural Networks (Anderson & Rosenfeld, eds., MIT Press, 2000) via Yuxi Liu's 2023 essay — a secondary relay, not this session's own direct read (direct.mit.edu, archive.org, and a Brown University repository page all failed to serve the book this session) — so it remains unpromotable per the vault's quote-provenance rule. Separately, an Internet Archive catalog record documents the "Neural Networks for Computing" conference series at Snowbird, Utah beginning in 1986, not 1985 — a discrepancy against the specific year this note and Wikipedia both carry, left unresolved rather than silently corrected. Net effect: the mechanism claim above (hard-limiting quantizers blocked hidden-layer training) stays primary-verified and unaffected; the specific "1985 Snowbird" biographical detail is now confirmed to rest on Wikipedia alone among sources checked to date, one tier below this note's own floor for a specific date/venue. See claim-widrow-1966-bootstrap-learning-preliminary-multilayer-attempt and claim-steinbuch-widrow-1965-comparison-not-multilayer-training for the rest of this session's findings, and the capture at10-inbox/raw/2026-09-05-do-steinbuch-widrow-1965-and-widrows-1966-bootstrap.md.
Source
“In 1987, Widrow, Winter, and Baxter looked back at the original Madaline I algorithm with the goal of developing a new technique that could adapt multiple layers of adaptive elements using the simpler hard-limiting quantizers.”