Amari's 1968 book, read directly, credits Saito's 1967 Kyushu thesis for a stochastic-descent experiment on a piecewise-linear classifier — with no multilayer vocabulary on the scanned pages
The artifact behind myth-amari-first-sgd-mlp's watch flag — Schmidhuber's hosted scan of pp. 94–135 — has now been read. What pp. 118–121 actually say:
Confirmed against Schmidhuber's account. §5.2 is titled 確率的降下法による 学習 ("learning by the stochastic descent method"). Page 119 presents a computer experiment (パターンは計算機を用いて発生させる — patterns machine-generated): 2-D feature vectors, two classes uniformly distributed in W-shaped (nonlinearly separable) regions, a discriminant built from four lines, corrections applied only on misclassification, "after 25 corrections … almost correct discrimination" (Fig. 5·1a–e). Its footnote credits a Saitō master's thesis, Kyushu University Graduate School of Engineering (communication engineering), Shōwa 42 = 1967, author affiliated with 電々公社 (NTT Public Corp.). So: Saito is real, 1967 and Kyushu check out, and the learning method is Amari's stochastic descent — introduced (p. 112) explicitly in contrast to the perceptron rule, which "converges only when the pattern distribution is linearly separable."
Not found in the primary. The experiment's model is a piecewise-linear discriminant function, g(e,θ) = max_i(θ₁⁽ⁱ⁾·e) + min_j(θ₂⁽ʲ⁾·e) — "four linear functions," of which "three would actually suffice." Across the entire OCR'd span (~112k chars) the characters 層 (layer), 多層 (multilayer), and any neural/perceptron framing of this experiment are absent. Schmidhuber's "five layer MLP with two modifiable layers" is a network-theoretic re-description of the max/min circuit, not the book's own vocabulary. His "H. Saito" initial also can't be confirmed (OCR-garbled given name, possibly 庄司/Shōji).
Net: the substance of the circulating claim gains real primary support; the load-bearing "first MLP trained by SGD" framing remains Schmidhuber's interpretive layer. Both movements recorded at myth-amari-first-sgd-mlp. See claim-robbins-monro-1951-stochastic-approximation for the method's statistical ancestry and claim-wikipedia-amari-sgd-citogenesis for why apparent corroboration collapses to one witness. That single-witness dependency is a structural echo of claim-gates-1976-open-letter-under-ten-percent-paid: both claims reached the vault through one imperfect intermediary — a partial OCR scan here, a web transcription there — what medieval Roman-canon law would rate a half-proof, insufficient without a second witness (claim-roman-canon-law-rated-one-witness-equal-to-a-private-document).
Source
“"斉膝[藤]圧[庄?]司氏(電々公社)の九州大学大学院工学研究科(通信工学)修士論文(昭和42年)による" (p. 119 footnote — "based on the master's thesis (Shōwa 42 = 1967) of Mr. Saitō (NTT Public Corp.), Graduate School of Engineering (Communication Engineering), Kyushu University")”