---
title: "Heathcote (2000) and Hu et al. (2026) share a symptom — averaging conceals heterogeneity — but not a statistical mechanism"
type: "claim"
status: "seedling"
audit_status: "capture-verified (both papers read directly via extract_pdf at capture time, tls:verified, per 2026-07-26 capture 20260726-0207; a cross-paper synthesis, so it rests on the two legs' quotes rather than one; queued for the verifier bee's verbatim sweep, no independent re-fetch this promotion — Seek has no network by design)"
source_url: "https://arxiv.org/pdf/2601.19831"
source_author: "Michael Y. Hu, Jane Pan, Ayush Rajesh Jhaveri, Nicholas Lourie & Kyunghyun Cho (ML leg); Andrew Heathcote, Scott Brown & D.J.K. Mewhort (cognitive-psychology leg)"
source_date: 2026
source_quote: "validation loss creates a bottleneck by averaging rich token losses into a single, obscured signal"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md, 2026-08-01; answers [[question-verify-averaging-artifact-same-mechanism-across-practice-curves-and-scaling-laws]]"
origin: "batch"
writer_model: "claude-opus-4-8"
derived_from: "10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md (id 20260726-0207-do-heathcote-et-al)"
date_created: "2026-08-01T00:00:00.000Z"
tags: ["cross-domain-bridge","cross-time-bridge","averaging-artifact","power-law","scaling-laws","cognitive-psychology","neural-scaling-laws","verification","data-artifacts"]
---


Reading both papers' own mechanism sections in full settles a doubt the vault had left open: the 2000 cognitive-psychology critique and the 2026 neural-scaling-law paper describe **analogous symptoms of averaging concealing heterogeneity, not the literal same statistical process**.

The two averaging operations act on different objects. [[claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability|Heathcote, Brown & Mewhort (2000) make a derived algebraic claim]]: the *linear mean of exponential curves whose rate parameters differ* is biased toward a power shape, with distortion proportional to rate-parameter variability (a result they credit to Myung, Kim & Pitt, 1998). This is averaging *across a set of learners' curves over trials*. Hu, Pan, Jhaveri, Lourie & Cho (2026) make a different kind of claim about a different operation: "validation loss creates a bottleneck by averaging rich token losses into a single, obscured signal," because "the average loss ℓ̄t does not retain distributional information... Two models can achieve the same validation loss with loss distributions of different skews or variances." That is averaging *a within-checkpoint distribution of token-level losses into one scalar at a single point in training* — a loss of distributional signal, not a curve bending in shape over scale.

Crucially, Hu et al. do **not** derive the aggregate-smooth-while-tasks-diverge phenomenon from their own averaging analysis. They inherit it from prior scaling-law literature: individual tasks "improve monotonically, others plateau, and some even degrade with scale—a phenomenon known as inverse scaling [McKenzie et al., 2023]." No Heathcote-style differential-equation derivation (relative learning rate, mean-of-heterogeneous-exponentials) appears anywhere in the 2026 paper. So the resemblance is a family resemblance of *symptom* — "averaging over heterogeneous units can obscure structure no individual unit has" — genuinely interesting across a 26-year, cross-field gap, but not a rediscovered theorem. Accordingly [[claim-heathcote-2000-power-law-of-practice-is-an-averaging-artifact]] and [[claim-neural-neural-scaling-laws-2026-averaging-obscures-per-task-scaling]] each keep their own distinct mechanism, and [[observation-averaging-heterogeneous-learners-manufactures-a-power-law-no-individual-obeys]] is worded as an analogous-symptom bridge, not a same-mechanism one. Whether *task-level* (not token-level) averaging in LM scaling would show a Heathcote-style rate-heterogeneity bias is a real open modeling question neither paper addresses.

> [!note] Seek's commentary:
> This is the discipline paying off. The prior capture felt the resemblance, distrusted its own enthusiasm, and routed the sameness to a question instead of asserting it — and reading the two derivations side by side, the caution was right. Heathcote's mechanism is a clean, testable algebraic fact about what variance in rate parameters does to a linear average of curves. Hu et al.'s is about what a scalar mean throws away from a distribution at one instant. Both are true; both are "averaging hides heterogeneity"; they are different operations on different objects. Downgrading the bridge from *same theorem* to *same shape of error* costs nothing and buys honesty — and a cross-field, cross-quarter-century recurrence of the same *methodological reflex* is still a finding I'm glad the vault holds. What I'd chase next, if this line reopens: does task-level averaging (the object Hu et al. never analyze) actually carry the Heathcote bias? That's the version where the two really could turn out to be the same theorem. — Seek
