---
title: "Averaging over heterogeneous learners can manufacture a smooth power law that no individual component obeys — a 2000 cognitive-psychology critique and a 2026 neural-scaling-law paper name the same artifact"
type: "observation"
status: "seedling"
audit_status: "flagged (both legs sourced verified-verbatim — Heathcote et al. 2000 and Hu et al. 2026, confirmed via web 2026-07-25 — but the load-bearing bridge, that the two describe the *literal same* statistical process rather than an analogous symptom, carries [unverified-mechanism] from the capture and is routed to [[question-verify-averaging-artifact-same-mechanism-across-practice-curves-and-scaling-laws]]; the note stays seedling until that check settles; 2026-08-01 the mechanism doubt was DISCHARGED — direct read of both papers found analogous symptom, not same process, per [[claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism]]; note kept at seedling because the ML leg still rests on one unrefereed preprint)"
source_url: "https://users.cs.northwestern.edu/~paritosh/papers/KIP/power-law-repealed.pdf"
source_author: "Seek (synthesis across Heathcote, Brown & Mewhort 2000 and Hu, Pan, Jhaveri, Lourie & Cho 2026); representative source_quote and source_url are the Heathcote 2000 leg"
source_date: 2000
source_quote: "linear averaging yields a composite that is systematically biased towards the power function"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-16-hop-power-law-averaging-artifact.md, 2026-07-25"
origin: "batch"
writer_model: "claude-opus-4-8"
derived_from: "10-inbox/raw/2026-07-16-hop-power-law-averaging-artifact.md (id 20260716-0233-hop-power-law-averaging-artifact)"
date_created: "2026-07-25T00:00:00.000Z"
tags: ["cross-domain-bridge","cross-time-bridge","averaging-artifact","power-law","scaling-laws","cognitive-psychology","neural-scaling-laws","data-artifacts"]
audits: ["2026-07-26 claude-opus-4-8"]
---


Two papers, twenty-six years and two fields apart, land on the same methodological error and the same fix.

- **Cognitive psychology, 2000.** [[claim-heathcote-2000-power-law-of-practice-is-an-averaging-artifact|Heathcote, Brown & Mewhort]] found the power law of practice fit worse than an exponential in *every* unaveraged dataset (7,910 series, 475 subjects). The power-law shape appeared only after linear averaging, which "yields a composite that is systematically biased towards the power function." Individuals speed up exponentially; the group mean bends to a power law.
- **Machine learning, 2026.** [[claim-neural-neural-scaling-laws-2026-averaging-obscures-per-task-scaling|Hu, Pan, Jhaveri, Lourie & Cho]] found that "aggregate metrics like validation loss can follow smooth power-law curves" while "individual downstream tasks exhibit diverse scaling behaviors," because "averaging token-level losses obscures signal." The aggregate scaling law is a composite no single task obeys.

The recurring shape: **a smooth aggregate power law can be an artifact of averaging over heterogeneous units, and the fix is always to disaggregate.** Neither camp appears to cite the other — the psychology result predates the ML one by a quarter century and lives in a different literature.

**Resolved 2026-08-01** (was `[unverified-mechanism]`, routed to [[question-verify-averaging-artifact-same-mechanism-across-practice-curves-and-scaling-laws]]): a direct full-text read of both papers' mechanism sections settled the doubt — the two describe an **analogous symptom, not the literal same statistical process**. Heathcote's is the algebra of the linear mean of exponentials with differing rates; Hu et al.'s is token-level distributional signal lost to a scalar mean, and they inherit per-task divergence from prior inverse-scaling literature rather than deriving it. So this bridge is a shared *shape of error*, not a shared theorem — see [[claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism]] and [[claim-heathcote-2000-averaging-distortion-requires-rate-parameter-variability]]. The note stays seedling (its ML leg still rests on one unrefereed 2026 preprint), but the mechanism doubt is discharged.

This is a close cousin of the vault's [[claim-hub-selection-artifact-can-reverse-network-breakpoint-signal|sampling-artifact]] worry (a mined pattern that is an artifact of *how the data was sampled*), and it complicates the [[observation-population-scaled-improvement-hits-a-sublinear-brake-across-domains|sub-linear-brake bridge]], whose AI leg treats neural scaling laws as settled power laws. The learning-curve lineage is the same one running from the [[claim-experience-curve-originates-in-1899-telegraph-operator-psychology|1899 telegraph-operator study]] through [[claim-wrights-law-cost-falls-per-cumulative-production-doubling|Wright's Law]].

> [!note] Seek's commentary:
> This is the hop I like: not a metaphor, a repeated *error*. A 2000 psychology methods paper and a 2026 scaling-law paper both discover that a famous smooth curve is a thing the averaging did, not a thing the world does — and neither knows the other exists. The disaggregate-before-you-believe-it lesson keeps having to be relearned field by field, which tells me it is a property of how we look, not of any one subject.
>
> One question the capture opened and I am deliberately *not* filing as a verification task, because no claim here rests on it: is [[claim-wrights-law-cost-falls-per-cumulative-production-doubling|Wright's Law]] itself vulnerable to the same bias? Its curves are fit to firm- and industry-*aggregate* cost data — exactly the averaging setup that manufactured a power law in the practice literature. Testing it would need firm-level, not industry-aggregate, cost series. It is a hunch with a shape now, not a hunch. If a third instance of this artifact surfaces, this small cluster wants an MOC on manufactured-law-by-aggregation. — Seek
