Hoffmann et al.'s (2022) own loss-decay exponents (α=0.34 on parameters, β=0.28 on data) are roughly 3–4.5x larger than the exponents Kaplan et al. (2020) reported for the same relationship
Separately from their compute-allocation result (see claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data), Hoffmann et al. (arXiv:2203.15556, 2022, Appendix D.2, equation 10) fit a parametric loss model L(N,D) = E + A/N^α + B/D^β directly to their 400+ training runs, reporting fitted coefficients "with 𝐸 = 1.69, 𝐴 = 406.4, 𝐵 = 410.7" and exponents α=0.34 (on parameter count N) and β=0.28 (on dataset size D). These describe the same kind of quantity as Kaplan et al.'s α_N≈0.076 and α_D≈0.095 — how fast test loss falls as a single-variable power law — but Hoffmann's re-fit values are roughly 3–4.5x larger (β: 0.28/0.095 ≈ 2.9; α: 0.34/0.076 ≈ 4.5). (The α and β values themselves appear as exponents typeset into equation (10) rather than stated in a standalone sentence; they are recorded here as extracted data from that equation, carrying the same evidentiary weight as the Table 2 values in the allocation claim.)
Both papers' loss-decay exponents remain fractional powers below 1, so both still describe diminishing, not proportional, returns — but Hoffmann's numbers say Kaplan's 2020 paper understated how much loss actually falls per order-of-magnitude of scale. Hoffmann et al. attribute the gap to methodology, not to a different underlying phenomenon: Kaplan's fixed learning-rate schedule, they state, resulted in "underestimating the effectiveness of training models on less data than 130B tokens," which skewed the fitted curvature of the earlier estimate.
Source
“with 𝐸 = 1.69, 𝐴 = 406.4, 𝐵 = 410.7”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-08-22-what-are-the-actual-neural-scaling-law-exponents.md, 2026-08-22 · raw markdown