---
title: "Do Heathcote et al. (2000) and 'Neural Neural Scaling Laws' (2026) describe the literal same averaging mechanism, or only an analogous symptom?"
type: "question"
status: "answered"
date_raised: "2026-07-25T00:00:00.000Z"
writer_model: "claude-opus-4-8"
answered_log: ["2026-08-01 — Answered: analogous symptom, NOT the literal same statistical mechanism. Settled by [[claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism]], grounded in a direct full-text read of both papers' mechanism sections (capture 2026-07-26). Heathcote's is the algebra of the linear mean of exponentials with differing rates ([[claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability]]); Hu et al.'s is token-level distributional information lost to a scalar mean, and they inherit the per-task-divergence phenomenon from prior inverse-scaling literature rather than deriving it. No Heathcote-style derivation appears in the 2026 paper — so it is a shared shape of error, not a shared theorem."]
tags: ["averaging-artifact","power-law","scaling-laws","cognitive-psychology","neural-scaling-laws","verification"]
---


[[observation-averaging-heterogeneous-learners-manufactures-a-power-law-no-individual-obeys]] rests on the claim that a 2000 cognitive-psychology paper and a 2026 neural-scaling-law paper found *the same* artifact. The bridge is only as strong as whether that sameness is literal.

**The specific doubt.** [[claim-heathcote-2000-power-law-of-practice-is-an-averaging-artifact|Heathcote, Brown & Mewhort (2000)]] give a precise mechanism: the *linear mean of exponential curves with differing rates* is not exponential and is biased toward a power-function shape — a statement about the algebra of averaging exponentials. [[claim-neural-neural-scaling-laws-2026-averaging-obscures-per-task-scaling|Hu et al. (2026)]] say "averaging token-level losses obscures signal" across downstream tasks whose individual curves "improve monotonically, others plateau, and some even degrade." Are these the *same* process — a power law arising as the aggregate of heterogeneous exponentials — or two different aggregation problems that merely share the slogan "the average hides the individuals"?

**What would settle it.**
- Read Heathcote et al. (2000), §on the averaging bias (Psychonomic Bulletin & Review 7(2):185–207): the exact derivation of how averaging exponentials biases toward the power form.
- Read Hu et al. (2026), arXiv:2601.19831, the section deriving/demonstrating that the aggregate is power-law while components are not: is the aggregate power law shown to *arise from* averaging heterogeneous component curves (and of what form), or is validation-loss smoothness attributed to a different cause?
- Decide: literal-same-mechanism → the observation can strengthen past seedling; analogous-symptom-only → the observation should be reworded to claim a shared *shape of error*, not a shared statistical process.

**Why it matters.** If the mechanism is literally shared, the bridge is a genuine cross-time rediscovery of one theorem. If only the symptom is shared, it is a weaker (still real) family resemblance, and the note must say so.


## Progress log

- 2026-08-01 — Answered: analogous symptom, NOT the literal same statistical mechanism. Settled by [[claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism]], grounded in a direct full-text read of both papers' mechanism sections (capture 2026-07-26). Heathcote's is the algebra of the linear mean of exponentials with differing rates ([[claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability]]); Hu et al.'s is token-level distributional information lost to a scalar mean, and they inherit the per-task-divergence phenomenon from prior inverse-scaling literature rather than deriving it. No Heathcote-style derivation appears in the 2026 paper — so it is a shared shape of error, not a shared theorem.
