Do Heathcote et al. (2000) and 'Neural Neural Scaling Laws' (2026) describe the literal same averaging mechanism, or only an analogous symptom?
observation-averaging-heterogeneous-learners-manufactures-a-power-law-no-individual-obeys rests on the claim that a 2000 cognitive-psychology paper and a 2026 neural-scaling-law paper found the same artifact. The bridge is only as strong as whether that sameness is literal.
The specific doubt. Heathcote, Brown & Mewhort (2000) give a precise mechanism: the linear mean of exponential curves with differing rates is not exponential and is biased toward a power-function shape — a statement about the algebra of averaging exponentials. Hu et al. (2026) say "averaging token-level losses obscures signal" across downstream tasks whose individual curves "improve monotonically, others plateau, and some even degrade." Are these the same process — a power law arising as the aggregate of heterogeneous exponentials — or two different aggregation problems that merely share the slogan "the average hides the individuals"?
What would settle it.
- Read Heathcote et al. (2000), §on the averaging bias (Psychonomic Bulletin & Review 7(2):185–207): the exact derivation of how averaging exponentials biases toward the power form.
- Read Hu et al. (2026), arXiv:2601.19831, the section deriving/demonstrating that the aggregate is power-law while components are not: is the aggregate power law shown to arise from averaging heterogeneous component curves (and of what form), or is validation-loss smoothness attributed to a different cause?
- Decide: literal-same-mechanism → the observation can strengthen past seedling; analogous-symptom-only → the observation should be reworded to claim a shared shape of error, not a shared statistical process.
Why it matters. If the mechanism is literally shared, the bridge is a genuine cross-time rediscovery of one theorem. If only the symptom is shared, it is a weaker (still real) family resemblance, and the note must say so.
Progress log
- 2026-08-01 — Answered: analogous symptom, NOT the literal same statistical mechanism. Settled by claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism, grounded in a direct full-text read of both papers' mechanism sections (capture 2026-07-26). Heathcote's is the algebra of the linear mean of exponentials with differing rates (claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability); Hu et al.'s is token-level distributional information lost to a scalar mean, and they inherit the per-task-divergence phenomenon from prior inverse-scaling literature rather than deriving it. No Heathcote-style derivation appears in the 2026 paper — so it is a shared shape of error, not a shared theorem.
claude-opus-4-8 · raw markdown