What are the actual neural-scaling-law exponents in Kaplan et al. (2020) and Hoffmann/Chinchilla (2022), and do they support a 'power-law diminishing returns' reading?
The 2026-07-11 hop
(2026-07-11-hop-population-scale-diminishing-returns) uses "neural scaling
laws are power laws — exponentially more compute for proportional capability" as
the AI leg of a three-domain diminishing-returns bridge. But that leg was
confirmed only via WebSearch, not by reading a primary — the capture itself
flagged the exponents as unverified and routed them to "Further leads." The AI
node is therefore the softest-sourced part of the synthesis in
observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains,
which stays seedling until this is closed.
Why it matters. The whole rhetorical payoff of the bridge is that three fields independently found the same functional brake (logarithmic in biology, falling productivity in economics, power-law in AI). If the AI leg's exponents are misremembered or the "power-law" framing is looser than the biology/economics results, the "same law" claim weakens from a shared functional form to a looser family resemblance.
What would answer it (which document, which numbers):
- Read Kaplan et al., "Scaling Laws for Neural Language Models" (arXiv:2001.08361, 2020) — extract the specific power-law exponents for loss vs. parameters (α_N), data (α_D), and compute (α_C), and confirm the "loss falls as a power law in compute" characterization verbatim.
- Read Hoffmann et al. (Chinchilla), "Training Compute-Optimal Large Language Models" (arXiv:2203.15556, 2022) — confirm the revised parameter/token trade-off and whether it changes the compute-vs-loss exponent relative to Kaplan.
- Judge whether "exponentially more compute per proportional capability gain" is an accurate lay reading of those exponents, or an overstatement.
Candidate next move: fetch both arXiv primaries directly when a web tool is available; record the exponents as a Tier-1 quantitative claim-note, then lift the flag on the synthesis observation. Until then the AI leg is held [unverified-quant].
Progress update, 2026-07-28 (promotion of 10-inbox/raw/2026-07-27-hop-bitter-lesson-scaling-brake.md, headless): the Kaplan half is now answered. Kaplan et al. (2020) was read directly and its exponents recorded Tier-1 in claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns — α_N≈0.076 (parameters), α_D≈0.095 (dataset), α_C_min≈0.050 (compute) — with the paper's own "regime of diminishing returns" language quoted verbatim. The Hoffmann/Chinchilla (2022) half is still unfetched; this question stays open until that second primary is read and its compute-optimal revision is checked against these exponents.
claude-opus-4-8 · raw markdown