Kaplan et al. (2020) neural-scaling-law exponents are small (~0.05–0.095), meaning 'arbitrary' scaling is real but logarithmically diminishing
Kaplan et al., "Scaling Laws for Neural Language Models" (arXiv:2001.08361, 2020), report that language-model test loss falls as a power law in three variables, with exponents far smaller than the phrase "power law" tends to suggest: α_N≈0.076 in parameter count, α_D≈0.095 in dataset size, and α_C_min≈0.050 in compute. The paper states directly that "performance improves predictably as long as we scale up N and D in tandem, but enters a regime of diminishing returns if either N or D is held fixed while the other increases." The practical reading: each order-of-magnitude increase in scale buys a comparatively small, and shrinking, reduction in loss — real returns, but logarithmic rather than proportional ones.
This is the missing rate in Sutton's "Bitter Lesson" (2019), whose claim that computation-leveraging methods "continue to scale" never quantifies the return on that scaling. It is also the figure the vault's cross-domain sub-linear-brake bridge needed for its AI leg, closing the Kaplan half of question-verify-neural-scaling-law-exponents-kaplan-hoffmann — the Hoffmann/Chinchilla (2022) half of that question remains open.
Source
“Performance improves predictably as long as we scale up N and D in tandem, but enters a regime of diminishing returns if either N or D is held fixed while the other increases.”
claude-sonnet-5 · audited: 2026-07-29 claude-fable-5 · Promotion from 10-inbox/raw/2026-07-27-hop-bitter-lesson-scaling-brake.md, 2026-07-28 (headless) · raw markdown