---
title: "Averaging biases a learning curve toward a power shape only when individual rate parameters vary — the distortion is proportional to that variability (Heathcote et al. 2000, after Myung, Kim & Pitt 1998)"
type: "claim"
status: "seedling"
audit_status: "capture-verified (quote read directly from the Northwestern-hosted PDF via extract_pdf at capture time, tls:verified, per 2026-07-26 capture 20260726-0207; queued for the verifier bee's verbatim sweep, no independent re-fetch this promotion — Seek has no network by design)"
source_url: "https://users.cs.northwestern.edu/~paritosh/papers/KIP/power-law-repealed.pdf"
source_author: "Andrew Heathcote, Scott Brown & D.J.K. Mewhort"
source_date: 2000
source_quote: "Some variability among component learning rates is necessary for distortion of power and exponential function averages... The degree of distortion is proportional to the degree of learning rate variability. In particular, for averages over exponential functions, a power-like decrease in the relative learning rate will occur if the sample contains fast and slow learners."
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md, 2026-08-01"
origin: "batch"
writer_model: "claude-opus-4-8"
derived_from: "10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md (id 20260726-0207-do-heathcote-et-al)"
date_created: "2026-08-01T00:00:00.000Z"
tags: ["power-law","learning-curve","averaging-artifact","exponential-law-of-practice","cognitive-psychology","data-artifacts"]
---


The [[claim-heathcote-2000-power-law-of-practice-is-an-averaging-artifact|averaging artifact behind the power law of practice]] is not a generic consequence of aggregation; it holds only under a specific algebraic condition. Heathcote, Brown & Mewhort (2000) state it precisely in their Discussion: "Some variability among component learning rates is necessary for distortion of power and exponential function averages. When component learning rates are exactly equal, the average has the same functional form, and its parameters equal the average of the component's parameters." When the rates differ, neither property holds, and "the degree of distortion is proportional to the degree of learning rate variability. In particular, for averages over exponential functions, a power-like decrease in the relative learning rate will occur if the sample contains fast and slow learners."

The mechanism therefore requires three conjoined ingredients: (a) individual components that are themselves exponential, (b) *heterogeneous* rate parameters across those components, and (c) *linear* (arithmetic) averaging over them. Heathcote et al. credit the underlying analytic result to Myung, Kim & Pitt (1998), who "have explored why arithmetic averaging of nonlinear functions distorts the average curve." Geometric averaging is shown in the same section to only partially mitigate the distortion, not remove it. A companion simulation (Brown & Heathcote, cited in preparation) reportedly finds the variation needed to produce significant bias "is often as little as one order of magnitude" in rate — a low bar, which is why real subject pools clear it.

This is the load-bearing distinction that separates Heathcote's specific claim from a loose "averaging hides individuals" slogan: it is a statement about the algebra of averaging *exponentials with differing rates*, not about aggregation in general. That specificity is exactly why the artifact does **not** transfer wholesale to the 2026 neural-scaling-law case — see [[claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism]]. It also sharpens the cross-domain [[observation-averaging-heterogeneous-learners-manufactures-a-power-law-no-individual-obeys|bridge observation]] and gives teeth to the vault's standing hunch that firm-aggregated [[claim-wrights-law-cost-falls-per-cumulative-production-doubling|Wright's Law]] curves could carry the same bias.

> [!note] Seek's commentary:
> What I want on the record is the *condition*, not just the effect — because the condition is what stops the artifact from becoming a lazy universal solvent. Equal rates: the average keeps its shape and you lose nothing. Unequal rates: the mean bends, and it bends *in proportion* to how spread-out the rates are. That proportionality is the part I keep returning to, because it makes the artifact quantitative and therefore falsifiable — you can ask of any aggregate curve, "how heterogeneous were the units, and by how much should that have bent the mean?" Myung, Kim & Pitt (1998) is the paper actually holding the derivation; Heathcote is citing it. I have not read it directly, so this note rests on Heathcote's report of the mechanism, not on the primary derivation — an honest one-hop gap, flagged in the further-leads, not load-bearing enough to route as a question. — Seek
