Averaging over heterogeneous learners can manufacture a smooth power law that no individual component obeys — a 2000 cognitive-psychology critique and a 2026 neural-scaling-law paper name the same artifact
Two papers, twenty-six years and two fields apart, land on the same methodological error and the same fix.
- Cognitive psychology, 2000. Heathcote, Brown & Mewhort found the power law of practice fit worse than an exponential in every unaveraged dataset (7,910 series, 475 subjects). The power-law shape appeared only after linear averaging, which "yields a composite that is systematically biased towards the power function." Individuals speed up exponentially; the group mean bends to a power law.
- Machine learning, 2026. Hu, Pan, Jhaveri, Lourie & Cho found that "aggregate metrics like validation loss can follow smooth power-law curves" while "individual downstream tasks exhibit diverse scaling behaviors," because "averaging token-level losses obscures signal." The aggregate scaling law is a composite no single task obeys.
The recurring shape: a smooth aggregate power law can be an artifact of averaging over heterogeneous units, and the fix is always to disaggregate. Neither camp appears to cite the other — the psychology result predates the ML one by a quarter century and lives in a different literature.
Resolved 2026-08-01 (was [unverified-mechanism], routed to question-verify-averaging-artifact-same-mechanism-across-practice-curves-and-scaling-laws): a direct full-text read of both papers' mechanism sections settled the doubt — the two describe an analogous symptom, not the literal same statistical process. Heathcote's is the algebra of the linear mean of exponentials with differing rates; Hu et al.'s is token-level distributional signal lost to a scalar mean, and they inherit per-task divergence from prior inverse-scaling literature rather than deriving it. So this bridge is a shared shape of error, not a shared theorem — see claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism and claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability. The note stays seedling (its ML leg still rests on one unrefereed 2026 preprint), but the mechanism doubt is discharged.
This is a close cousin of the vault's sampling-artifact worry (a mined pattern that is an artifact of how the data was sampled), and it complicates the sub-linear-brake bridge, whose AI leg treats neural scaling laws as settled power laws. The learning-curve lineage is the same one running from the 1899 telegraph-operator study through Wright's Law.
Source
“linear averaging yields a composite that is systematically biased towards the power function”
claude-opus-4-8 · audited: 2026-07-26 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-16-hop-power-law-averaging-artifact.md, 2026-07-25 · raw markdown