Do Steinbuch & Widrow 1965 and Widrow's 1966 'Bootstrap Learning' actually document the abandoned multilayer-training attempt — and what does Widrow's oral history say about the 1985 return?
verifies: question-verify-widrow-1965-66-papers-multilayer-attempt
Answers question-verify-widrow-1965-66-papers-multilayer-attempt, raised during promotion of claim-widrow-abandoned-multilayer-training-until-1985-backprop and claim-madaline-rule-i-first-layer-trainable-second-fixed. Both 1965–66 papers named in the question were fetched directly via extract_pdf and read in full (Steinbuch & Widrow 1965 via pdftotext, both tls: verified; Widrow 1966 via OCR — a 40-page scan of a conference reprint, method ocr, tls: verified). The ETHW oral history, previously logged elsewhere in this vault as a known-blocked route, resolved this session via archive_page and was read in full. No recognition signals per the safety spec fired on any page fetched this session (Stanford's own ISL site, ETHW, Wikipedia, Yuxi Liu's personal blog, and archive.org's catalog page all read as ordinary academic/reference content — no addressed-to-AI language, override language, claimed authority, or urgency framing).
Claim: Steinbuch & Widrow's 1965 paper is a structural and capacity comparison between two existing single-decision-layer classifiers, not a study of multilayer network training
verifies: question-verify-widrow-1965-66-papers-multilayer-attempt
Claim type: historical/definitional (what a specific named document is about). Sourcing floor: Tier 3–4 acceptable; met here at Tier 1 by direct full read.
The paper, a two-page "Short Note" in IEEE Transactions on Electronic Computers, resulted from a May 1964 visit by Karlsruhe's Karl Steinbuch to Stanford and compares Steinbuch's "Learning Matrix" against Widrow's "Madaline" on number of outputs, crosspoint weights, adaptations required for training, and equipment ratios (Table I, Table II, Table III in the source). Both systems compared are single-decision architectures: the Learning Matrix assigns one output line per pattern class trained in one pass, and the Madaline described here is the one-Adaline-per-decision or majority/OR-combined form Widrow and Hoff had already published training rules for in 1960–63. The paper contains no discussion of hidden layers, credit assignment across layers, or training algorithms for networks with more than one adaptive stage — the word "multilayer" does not appear. The date-coincidence the original research question flagged (this paper falls in the 1965–66 window named for "the" abandoned multilayer attempt) is resolved: this specific paper is not that document. It is a bibliometric/capacity comparison paper unrelated to the multilayer-training problem, despite carrying Widrow's name and 1965 date.
Claim: Widrow's 1966 "Bootstrap Learning" paper does describe a genuine, preliminary, non-gradient attempt at training multilayer networks of adaptive threshold elements — a separate and independently primary-documented instance of the 1960s multilayer-training effort
verifies: question-verify-widrow-1965-66-papers-multilayer-attempt
Claim type: historical/technical-mechanism (what a specific attempt at multilayer training consisted of). Sourcing floor: Tier 1–2 required for the mechanism; met by direct primary read.
Unlike the 1965 note, this paper is squarely on-topic. Its introduction states its subject includes "convergent adaptation procedures for multilayered and more generally-connected networks of adaptive threshold elements," and its closing "Current and Future Research" section reports: "Preliminary studies have been made with some success toward the development of adaptation algorithms for multilayered networks of adaptive threshold elements using the selective bootstrap principle." The mechanism described is a global, non-gradient reinforcement scheme, distinct from the Madaline Rule I architecture (adaptive first layer, fixed second layer) already documented in claim-madaline-rule-i-first-layer-trainable-second-fixed: "If performance observed at a set of output terminals is 'better than average,' every element in the net receives positive bootstrap adaptation. If output-terminal performance is poorer than average, then all elements receive negative bootstrap adaptation" — every adaptive element in the network is nudged the same direction based only on aggregate output quality, with no per-element credit assignment. The paper frames this network-wide extension as future/ongoing work ("It is conjectured that formula 37 will be usable in predicting the rate of adaptation in such networks") rather than a completed result; the bulk of the paper analytically derives and experimentally verifies the single-element case (applied to simulated Blackjack play). This is corroborated independently by Widrow & Lehr's own 1990 retrospective, which states Widrow "devised a reinforcement learning algorithm called 'punish/reward' or 'bootstrapping'" and, separately, that the group's work "switched to adaptive filtering and adaptive signal processing" only "after attempts t[o] develop learning rules for networks with multiple adaptive layers were unsuccessful" (two-column PDF scan; these are adjacent clauses of one sentence in the original, confirmed present by direct read but not quotable as one unbroken string due to the column-merge artifact, per the convention already used in claim-madaline-rule-i-first-layer-trainable-second-fixed). Together, the two 1965–66 papers resolve cleanly: the 1965 note is unrelated to multilayer training; the 1966 paper is a real, primary-documented, if preliminary and reinforcement-based rather than gradient-based, multilayer-training attempt — a second and earlier-dated attempt than the Madaline Rule I framing this vault already carries.
Claim: Widrow's directly-read 1997 oral-history transcript contains no mention of a 1985 Snowbird, Utah conference, no return-to-neural-networks narrative, and none of the "sharp quantizers / sigmoids / no one knew" language commonly attributed to him — this specific primary document does not carry the story built on top of it
verifies: question-verify-widrow-1965-66-papers-multilayer-attempt
Claim type: historical (what a specific primary document does and does not contain). Sourcing floor: Tier 3–4 acceptable for an absence-in-a-document claim; met here at Tier 1 by a complete direct read of the full transcript (both halves, offset 1–402 and 403–526, ending "Retrieved from... Well, thank you very much for the interview.").
The full 1997 interview (conducted by Andrew Goldstein for the IEEE History Center, 14 March 1997, Stanford) covers Widrow's childhood, MIT education, quantization-noise PhD thesis, arrival at Stanford, the Adaline/Madaline/LMS work, an anecdote about training speed against Rosenblatt's perceptron, adaptive filtering, telephone equalization, and adaptive antennas — and ends there. It does not mention 1985, Snowbird, backpropagation, sigmoids, or a return to neural-network research at all. The interview's own table of contents (14 sections) confirms this is the complete transcript, not a truncated fetch. This is a directly checkable negative finding: the "1985 Snowbird" story — repeated in Wikipedia's Bernard Widrow article and in this vault's own claim-widrow-abandoned-multilayer-training-until-1985-backprop (there flagged as resting on "Widrow's (unread) oral history") — does not trace to this particular oral-history document, now that it has been read.
Claim: The specific "sharp quantizers / need sigmoids / no one knew anything about it" quote does exist, but traces to a different oral history (Talking Nets, 2000) than the ETHW transcript, and neither source available for direct or indirect check pairs that quote with "1985" or "Snowbird" — the specific date/venue rests only on a Tier 4 source and conflicts with the documented Snowbird conference's actual first year
verifies: question-verify-widrow-1965-66-papers-multilayer-attempt
Claim type: quantitative (a specific date and venue) plus historical. Sourcing floor: Tier 1–2 required for the date/venue; not met — flagged accordingly.
Wikipedia's Bernard Widrow article states: "At a 1985 conference in Snowbird, Utah, he noticed that neural network research was returning, and he also learned of the backpropagation algorithm." Wikipedia is Tier 4 and cites this passage only via a blended footnote cluster (the ETHW oral history, Talking Nets: An Oral History of Neural Networks [Anderson & Rosenfeld, eds., MIT Press, 2000], and Magoun's 2014 Proceedings of the IEEE "A Nonrandom Walk Down Memory Lane With Bernard Widrow"). The ETHW transcript, read directly above, does not contain it. Talking Nets itself could not be fetched directly this session (direct.mit.edu returned HTTP 403 to both WebFetch and archive_page; archive.org's copy is access-restricted; Brown University's "Talking Nets Archive" repository page bot-blocked the fetch), and the Magoun 2014 IEEE piece likewise could not be retrieved (IEEE Xplore's PDF endpoint did not serve a PDF; academia.edu returned HTTP 403). The closely-related quote — "Backprop would not work with the kind of neurons that we were using because the neurons we were using all had quantizers that were sharp. In order to make backprop work, you have to have sigmoids; you have to have a smooth nonlinearity … no one knew anything about it at that time. This was long before Paul Werbos. Backprop to me is almost miraculous." — is real and does appear in Yuxi Liu's 2023 essay, who attributes it in-line to "(Rosenfeld and Anderson 2000)," i.e. Talking Nets. Per the vault's quote-provenance rule, this quote is relayed through a summarizing secondary layer (Yuxi Liu's blog), not obtained by this session's own direct fetch of the primary book chapter, so it is recorded here as [unverified-quote — needs direct read] rather than promoted as Tier-1-grounded — and notably, even as relayed, it carries no date or venue at all; Yuxi Liu's surrounding text only says Widrow "heard about the 'miraculous' backpropagation in the 1980s." The specific "1985"/"Snowbird" pairing therefore currently rests on Wikipedia alone among the sources checked this session — marked [unverified-quant — needs primary]. Separately, an Internet Archive catalog record for the published proceedings of the "Neural Networks for Computing" conference series at Snowbird, Utah states plainly: "Proceedings of the Conference on Neural Networks for Computing held at Snowbird, Utah on April 13-16, 1986" — a documented first instance of that specific annual conference in 1986, not 1985. This is a discrepancy worth flagging rather than resolving: either Widrow's return predates the well-documented Snowbird conference series by a year, or the year has drifted in the retelling, or "the" 1985 conference is a different, undocumented gathering. This vault's own existing claim note's "1985 Snowbird" detail should be treated as unverified pending a direct read of Talking Nets or the Magoun 2014 piece.
Further leads
- Talking Nets: An Oral History of Neural Networks (Anderson & Rosenfeld, eds., MIT Press, 2000), chapter on Bernard Widrow — the most likely primary home of the "sharp quantizers/sigmoids" quote and any 1985/Snowbird detail; direct.mit.edu 403'd, archive.org copy access-restricted, Brown University repository bot-blocked this session. Worth a renewed attempt or manual consultation.
- Magoun, Alexander B., "A Nonrandom Walk Down Memory Lane With Bernard Widrow," Proceedings of the IEEE 102(10):1622–1629 (Oct. 2014) — an edited oral history drawing on IEEE History Center material; IEEE Xplore and academia.edu links both failed to serve content this session.
- Widrow, B., Cybernetics 2.0 (Springer Nature, 2023) — Widrow's own recent book, cited by Yuxi Liu for figures and a preface anecdote; not fetched, could carry a first-person "1985"/Snowbird account in his own words.
- The rest of Widrow & Lehr 1990 (Sections VI–VIII, ~1400 lines not read in this session) — covers Madaline Rule II/III and backpropagation mechanics in more depth; already partially cited in existing vault claims.
- Rosenblatt's own "back-propagating error-correction procedure" (Rosenblatt 1962, vol. 55, chap. 13) — per Yuxi Liu, an earlier and distinct use of the term "backpropagation" for a heuristic (non-gradient) two-layer perceptron training hack, predating both Widrow's bootstrap paper and Werbos — a candidate for its own claim on terminology priority versus mechanism priority.
Entity candidates
- Karl Steinbuch — person — the older, foundational figure the 1965 note's comparison is measured against: inventor of the "Learning Matrix" at Karlsruhe (from 1959), the system Widrow's Madaline is being benchmarked relative to in the one paper from this pair that turned out not to be about multilayer training.
- Marcian "Ted" Hoff — person — Widrow's co-inventor of the LMS algorithm and Madaline; already covered via claim-ted-hoff-widrow-phd-student-architected-intel-4004, but worth a direct entity page given his structural role across all three primary sources read this session.
- Andrew Goldstein — person — the IEEE History Center interviewer who conducted the 1997 Widrow oral history; a first encounter for this vault as a named oral-history interviewer, distinct from the interview subject.
- David Rumelhart — person — per Yuxi Liu's essay (itself drawing on his own Talking Nets interview), the connectionist who independently arrived at backpropagation via a "generalized delta rule" built explicitly on Widrow's ADALINE/delta rule — the direct causal link between the 1960s Widrow work captured here and the 1980s revival, worth its own claim in a future capture.
- Geoffrey Hinton — person — per the same essay, spent roughly two years actively arguing against and avoiding backpropagation before adopting it in 1985/86; a parallel "resistance before adoption" case to Widrow's own, and a candidate entity page if not already present.
- Talking Nets: An Oral History of Neural Networks (Anderson & Rosenfeld, eds., 2000) — concept/work — the primary oral-history volume this whole question ultimately routes back to; currently unreachable by this vault's tooling, worth flagging as a standing access gap alongside ethw.org.