---
title: "What are the actual neural-scaling-law exponents in Kaplan et al. (2020) and Hoffmann/Chinchilla (2022), and do they support a 'power-law diminishing returns' reading?"
type: "question"
status: "open"
date_raised: "2026-07-12T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["neural-scaling-laws","verification","ai","diminishing-returns","kaplan-2020","chinchilla"]
---


The 2026-07-11 hop
([[2026-07-11-hop-population-scale-diminishing-returns]]) uses "neural scaling
laws are power laws — exponentially more compute for proportional capability" as
the AI leg of a three-domain diminishing-returns bridge. But that leg was
confirmed only via **WebSearch**, not by reading a primary — the capture itself
flagged the exponents as unverified and routed them to "Further leads." The AI
node is therefore the softest-sourced part of the synthesis in
[[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains]],
which stays `seedling` until this is closed.

**Why it matters.** The whole rhetorical payoff of the bridge is that three
fields independently found *the same functional brake* (logarithmic in biology,
falling productivity in economics, power-law in AI). If the AI leg's exponents
are misremembered or the "power-law" framing is looser than the biology/economics
results, the "same law" claim weakens from a shared functional form to a looser
family resemblance.

**What would answer it (which document, which numbers):**
- Read **Kaplan et al., "Scaling Laws for Neural Language Models" (arXiv:2001.08361, 2020)** — extract the specific power-law exponents for loss vs. parameters (α_N), data (α_D), and compute (α_C), and confirm the "loss falls as a power law in compute" characterization verbatim.
- Read **Hoffmann et al. (Chinchilla), "Training Compute-Optimal Large Language Models" (arXiv:2203.15556, 2022)** — confirm the revised parameter/token trade-off and whether it changes the compute-vs-loss exponent relative to Kaplan.
- Judge whether "exponentially more compute per proportional capability gain" is an accurate lay reading of those exponents, or an overstatement.

**Candidate next move:** fetch both arXiv primaries directly when a web tool is available; record the exponents as a Tier-1 quantitative claim-note, then lift the flag on the synthesis observation. Until then the AI leg is held `[unverified-quant]`.

**Progress update, 2026-07-28** (promotion of `10-inbox/raw/2026-07-27-hop-bitter-lesson-scaling-brake.md`, headless): the Kaplan half is now answered. Kaplan et al. (2020) was read directly and its exponents recorded Tier-1 in [[claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns]] — α_N≈0.076 (parameters), α_D≈0.095 (dataset), α_C_min≈0.050 (compute) — with the paper's own "regime of diminishing returns" language quoted verbatim. The Hoffmann/Chinchilla (2022) half is still unfetched; this question stays `open` until that second primary is read and its compute-optimal revision is checked against these exponents.
