Heathcote (2000) and Hu et al. (2026) share a symptom — averaging conceals heterogeneity — but not a statistical mechanism
Reading both papers' own mechanism sections in full settles a doubt the vault had left open: the 2000 cognitive-psychology critique and the 2026 neural-scaling-law paper describe analogous symptoms of averaging concealing heterogeneity, not the literal same statistical process.
The two averaging operations act on different objects. Heathcote, Brown & Mewhort (2000) make a derived algebraic claim: the linear mean of exponential curves whose rate parameters differ is biased toward a power shape, with distortion proportional to rate-parameter variability (a result they credit to Myung, Kim & Pitt, 1998). This is averaging across a set of learners' curves over trials. Hu, Pan, Jhaveri, Lourie & Cho (2026) make a different kind of claim about a different operation: "validation loss creates a bottleneck by averaging rich token losses into a single, obscured signal," because "the average loss ℓ̄t does not retain distributional information... Two models can achieve the same validation loss with loss distributions of different skews or variances." That is averaging a within-checkpoint distribution of token-level losses into one scalar at a single point in training — a loss of distributional signal, not a curve bending in shape over scale.
Crucially, Hu et al. do not derive the aggregate-smooth-while-tasks-diverge phenomenon from their own averaging analysis. They inherit it from prior scaling-law literature: individual tasks "improve monotonically, others plateau, and some even degrade with scale—a phenomenon known as inverse scaling [McKenzie et al., 2023]." No Heathcote-style differential-equation derivation (relative learning rate, mean-of-heterogeneous-exponentials) appears anywhere in the 2026 paper. So the resemblance is a family resemblance of symptom — "averaging over heterogeneous units can obscure structure no individual unit has" — genuinely interesting across a 26-year, cross-field gap, but not a rediscovered theorem. Accordingly claim-heathcote-2000-power-law-of-practice-is-an-averaging-artifact and claim-neural-neural-scaling-laws-2026-averaging-obscures-per-task-scaling each keep their own distinct mechanism, and observation-averaging-heterogeneous-learners-manufactures-a-power-law-no-individual-obeys is worded as an analogous-symptom bridge, not a same-mechanism one. Whether task-level (not token-level) averaging in LM scaling would show a Heathcote-style rate-heterogeneity bias is a real open modeling question neither paper addresses.
Source
“validation loss creates a bottleneck by averaging rich token losses into a single, obscured signal”
claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md, 2026-08-01; answers [[question-verify-averaging-artifact-same-mechanism-across-practice-curves-and-scaling-laws]] · raw markdown