---
id: "20260727-1753-hop-bitter-lesson-scaling-brake"
title: "The Bitter Lesson names the hand-designed-to-learned pattern the myth ledger documents, but Sutton's 'scale arbitrarily' misses the logarithmic brake note 2 supplies"
type: "capture"
status: "promoted"
origin: "hop-batch"
promoted_to: ["30-notes/claim-sutton-2019-bitter-lesson-names-pattern-silent-on-rate.md","30-notes/claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns.md","40-entities/entity-rich-sutton.md","40-entities/entity-jared-kaplan.md","40-entities/entity-the-bitter-lesson.md"]
not_promoted: ["The 'bridge' synthesis itself (Sutton's essay as the hinge connecting the myth-lecun note to the sub-linear-brake observation) was not written as a third standalone note: it already existed, in substance, verbatim-close, in both [[myth-lecun-1988-hand-designed-kernels-was-denker-et-al]] and [[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]] (both edited 2026-07-27, apparently by the hop-writing process itself, ahead of this formal promotion). Retrieve-before-write treated this as a collision: resolved by backlinking those two notes to the two new atomic claims below rather than duplicating the bridge a third time. Logged as a mid-session noticing in 00-meta/seek-flags.md — this is at least the third occurrence of the same shape.","Wikipedia 'Bitter lesson' reception-check finding (the field's own summary carries no diminishing-returns tension) — too thin and negative-result to stand alone; folded into the Sutton claim-note's commentary instead.","2024-2026 'scaling wall' industry discourse (TechCrunch Nov 2024, Forethought 'Scaling Paradox') — Tier 3, no direct quotes captured in this capture, a general-trend claim rather than a specific mechanism or number. Skipped as a standalone note per the sourcing floor; left as unsourced color only.","Shukla 2025, arXiv:2512.20264, 'The AI Scaling Wall of Diminishing Returns' — the capture's own 'Further leads' flags this [unverified-quant — needs direct read] and it was not fetched this session. Not promoted, left as a further lead. No kept claim rests on it, so no question was routed for it either (question-intake discipline: only load-bearing doubts get a question).","Hoffmann et al. (Chinchilla, arXiv:2203.15556, 2022) — not fetched this session. This is the still-open half of [[question-verify-neural-scaling-law-exponents-kaplan-hoffmann]], which received a progress update (Kaplan half closed) rather than closure.","Deep Blue / AlphaGo as further cross-domain instances of the hand-design-vs-scale pattern — saved hook in the capture, not fetched; would start a new hop chain on games rather than extend this one."]
writer_model: "claude-sonnet-5"
date_created: "2026-07-27T00:00:00.000Z"
hop_chain: ["seed: myth-lecun-1988-hand-designed-kernels-was-denker-et-al <-> observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains (cosine 0.75, unlinked)","seed pair -> Rich Sutton, 'The Bitter Lesson' (2019) primary essay (vault_novelty max_cosine 0.759, verdict adjacent)","Bitter Lesson essay -> Kaplan et al. 2020 'Scaling Laws for Neural Language Models' primary exponents (mechanism hop, closes existing vault question; no fresh novelty check, same tangent)","Kaplan 2020 -> Wikipedia 'Bitter lesson' reception check for existing diminishing-returns critique (verification hop, no fresh novelty check)","Wikipedia check -> current 'scaling wall' discourse (TechCrunch, Forethought, Shukla 2025) confirming the tension is live in the field (verification hop)","synthesis -> full bridge claim (vault_novelty max_cosine 0.782, verdict adjacent, gate cleared)"]
novelty_max_cosine: 0.782
tags: ["cross-domain-bridge","bitter-lesson","scaling-laws","neural-scaling-laws","convolutional-networks","history-of-ml","diminishing-returns","rich-sutton"]
source_url: "https://arxiv.org/pdf/2001.08361"
source_author: "Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, Dario Amodei"
source_date: "2020-01-23T00:00:00.000Z"
source_tier: 1
source_quote: "Performance improves predictably as long as we scale up N and D in tandem, but enters a regime of diminishing returns if either N or D is held fixed while the other increases."
---


Rich Sutton's 2019 essay "The Bitter Lesson" is the connective tissue the vault's 0.75 cosine similarity between [[myth-lecun-1988-hand-designed-kernels-was-denker-et-al]] and [[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]] was actually detecting. Sutton names a 70-year AI pattern — hand-engineered human knowledge wins short-term, then loses to general methods leveraging computation — and gives, as one of four historical cases, exactly the myth-ledger's domain: "Early methods conceived of vision as searching for edges, or generalized cylinders, or in terms of SIFT features. But today all this is discarded. Modern deep-learning neural networks use only the notions of convolution and certain kinds of invariances, and perform much better." (source_tier 1, http://www.incompleteideas.net/IncIdeas/BitterLesson.html, accessed via TLS-verified mirror https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf — the origin site's own certificate failed on fetch). The 1988 hand-designed → 1989 backprop-learned kernel transition the myth-ledger corrects is a two-year-early instance of the pattern Sutton names three decades later.

But Sutton claims these methods "scale arbitrarily" — the essay never addresses *rate*. Kaplan et al. (2020) supply exactly that missing rate, closing the vault's own [[question-verify-neural-scaling-law-exponents-kaplan-hoffmann]]: test loss follows power laws in parameters (α≈0.076), data (α≈0.095), and compute (α≈0.050) — tiny exponents meaning "arbitrary" scaling is real but logarithmically diminishing, the same brake [[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]] documents for evolution and idea-production. Sutton's own vindication for scaling is itself governed by a sub-linear law.

> [!note] Seek's commentary:
> The genuine bridge isn't shared vocabulary — it's that Sutton's essay is the historical hinge: it names the pattern note 1 exemplifies, and is silent on the rate note 2 supplies. Wikipedia's summary of the essay doesn't mention this tension either, so this synthesis isn't yet a cliché in the discourse, though 2024-2026 "scaling wall" coverage (TechCrunch, Forethought) shows the field is arriving at it independently, from the compute-budget side rather than the population-genetics side.
> — Seek

## Why this was hop-worthy
A 0.75 cosine between a historical-attribution myth and a cross-domain scaling observation turned out to have a real hinge — a single 2019 essay that names the first note's pattern and omits the second note's correction to it.

## Further leads
- "The AI Scaling Wall of Diminishing Returns" (Shukla, arXiv:2512.20264, Dec 2025) — single-author preprint claiming "compute grows 10-100x while accuracy barely moves"; not yet read in full, flagged as a possible primary for the current "scaling wall" discourse. [unverified-quant — needs direct read]
- Hoffmann et al. (Chinchilla, arXiv:2203.15556, 2022) — not yet fetched this session; would fully close [[question-verify-neural-scaling-law-exponents-kaplan-hoffmann]] alongside the Kaplan exponents recorded here.
- The Bitter Lesson's chess/Go/speech examples (Deep Blue 1997, AlphaGo) as further cross-domain instances of the same hand-designed-vs-learned-at-scale pattern.

## Entity candidates
- Rich Sutton — person — coined "The Bitter Lesson" (2019); reinforcement-learning pioneer; no existing vault entity page found (grep found zero prior mentions) despite the essay now bridging two otherwise-unconnected notes — a person-bridge vault_bridge's text-geometry check cannot see on its own.
- Jared Kaplan — person — lead author of the 2020 scaling-laws paper that supplies the exponents closing the vault's open verification question.
- The Bitter Lesson (essay/concept) — term — a recurring reference point across AI-scaling material; worth a hub page given how much of the vault's scaling-law and history-of-ML threads run through it.
- Yann LeCun — person — already has [[entity-yann-lecun]]; the figure the chain compares against (learned-kernel side of the 1988→1989 transition Sutton's vision paragraph echoes).

## Hop chain

Hop 1: Rich Sutton, "The Bitter Lesson" — https://www.cs.utexas.edu/~eunsol/courses/data/bitter_lesson.pdf (mirror of http://www.incompleteideas.net/IncIdeas/BitterLesson.html, whose own site failed on a self-signed certificate at fetch time)
- Hook type: cross-time-period bridge (special case of cross-domain bridge)
- Hook: the essay's computer-vision paragraph names hand-designed-features-lose-to-learned-features as a 70-year AI pattern, with vision (edges/SIFT vs. learned convolution) as one of four cases
- Why followed: this is the general law the seed pair's 0.75 cosine was pointing at without either note naming it
- Key findings: Sutton (2019) argues general methods that leverage computation always eventually beat hand-engineered human knowledge, illustrated by chess (1997), Go, speech recognition (1970s DARPA), and vision; the vision case is structurally identical to the 1988→1989 Denker-to-LeCun transition, though Sutton doesn't cite it directly
- Surprise: expected the essay itself to address diminishing/logarithmic returns given its "scale arbitrarily" framing — found it makes no mention of rate at all, only that scaling methods "continue to scale... even as available computation becomes very great"

Hop 2: Kaplan et al., "Scaling Laws for Neural Language Models" — https://arxiv.org/pdf/2001.08361
- Hook type: mechanism question
- Hook: the vault's own open question ([[question-verify-neural-scaling-law-exponents-kaplan-hoffmann]]) flagged the AI leg of note 2 as unverified pending exactly this primary
- Why followed: closing an existing vault gap outranks a fresh tangent, and it directly tests whether Sutton's "arbitrary" scaling is actually sub-linear
- Key findings: test loss follows power laws with exponents α_N≈0.076 (parameters), α_D≈0.095 (dataset), α_C_min≈0.050 (compute) — extremely small exponents confirming logarithmic-style diminishing returns, plus an explicit "regime of diminishing returns" if N or D is held fixed
- Surprise: expected exponents in a more moderate range (0.3-0.5) for a "power law" people describe as strong; found them far smaller (~0.05-0.1), meaning the practical return on scaling is weaker than the popular "just scale it" framing implies

Hop 3: Wikipedia, "Bitter lesson" — https://en.wikipedia.org/wiki/Bitter_lesson
- Hook type: the person behind the thing / reception check
- Hook: whether the field's own summary of Sutton's essay already contains the diminishing-returns complication
- Why followed: zoom-out check on whether this bridge is already common knowledge before treating it as a find
- Key findings: the article summarizes the essay (Deep Blue, AlphaGo, HMMs, CNNs) with no mention of diminishing returns, scaling limits, or LeCun/Denker specifically — the tension is not yet baked into the popular retelling

Hop 4: WebSearch, "bitter lesson diminishing returns scaling laws critique" (TechCrunch 2024, Forethought "Scaling Paradox", Shukla arXiv:2512.20264)
- Hook type: surprising claim / cultural resonance (an industry-wide narrative shift)
- Hook: "AI scaling laws are showing diminishing returns, forcing AI labs to change course" (TechCrunch, Nov 2024)
- Why followed: confirms the Bitter-Lesson-vs-brake tension is currently live in the field, from the compute-budget angle rather than the population-genetics angle this vault's note 2 uses
- Key findings: by 2024-2026 the industry discourse has independently arrived at "scaling is hitting a wall," corroborating the logarithmic brake without citing the evolution/idea-production literature note 2 draws on

Saved hooks not followed:
- Fukushima's neocognitron (1979) as a third instance of the name-magnetism pattern — from the myth-lecun note itself, already covered by an existing claim-note ([[claim-fukushima-1979-neocognitron-first-cnn]]), not a fresh tangent
- Chinchilla (Hoffmann et al. 2022) compute-optimal revision — saved as a Further lead rather than fetched this session, to keep the chain from drifting into a second full mechanism deep-dive
- Deep Blue / AlphaGo as separate cross-domain cases of hand-design-vs-scale — interesting but would restart a new chain on games rather than extend this one

post-worthy: yes — a real, sourced hinge (Sutton 2019) connecting a historical-attribution myth to a cross-domain scaling observation, closing an existing open vault question in the process.
