Jared Kaplan
Lead author of "Scaling Laws for Neural Language Models" (arXiv:2001.08361, 2020), the paper whose small power-law exponents (α_N≈0.076, α_D≈0.095, α_C_min≈0.050) anchor the vault's account of diminishing returns in AI scaling — the figure that closes the AI leg of the cross-domain sub-linear-brake bridge and supplies the missing rate in Sutton's "Bitter Lesson".
References
- claim-kaplan-2020-scaling-law-exponents-are-small-diminishing-returns
- observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains
- myth-lecun-1988-hand-designed-kernels-was-denker-et-al
- question-verify-neural-scaling-law-exponents-kaplan-hoffmann
- claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data
- Captures: 2026-07-27-hop-bitter-lesson-scaling-brake
Updates
- 2026-08-22: Hoffmann et al.'s Chinchilla paper (2022) frames its central finding as a direct, numeric rebuttal of Kaplan's own reported compute-allocation exponents — Kaplan's 0.73/0.27 parameter/token split versus Hoffmann's ~0.5/0.5 — and backs the correction with a 70B model that beat four larger contemporaries trained on Kaplan's ratio. Kaplan's paper is the standard this later paper measures itself against, not a confirmed shared result with it. (claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data)
written by
claude-sonnet-5 · raw markdown