Averaging biases a learning curve toward a power shape only when individual rate parameters vary — the distortion is proportional to that variability (Heathcote et al. 2000, after Myung, Kim & Pitt 1998)
The averaging artifact behind the power law of practice is not a generic consequence of aggregation; it holds only under a specific algebraic condition. Heathcote, Brown & Mewhort (2000) state it precisely in their Discussion: "Some variability among component learning rates is necessary for distortion of power and exponential function averages. When component learning rates are exactly equal, the average has the same functional form, and its parameters equal the average of the component's parameters." When the rates differ, neither property holds, and "the degree of distortion is proportional to the degree of learning rate variability. In particular, for averages over exponential functions, a power-like decrease in the relative learning rate will occur if the sample contains fast and slow learners."
The mechanism therefore requires three conjoined ingredients: (a) individual components that are themselves exponential, (b) heterogeneous rate parameters across those components, and (c) linear (arithmetic) averaging over them. Heathcote et al. credit the underlying analytic result to Myung, Kim & Pitt (1998), who "have explored why arithmetic averaging of nonlinear functions distorts the average curve." Geometric averaging is shown in the same section to only partially mitigate the distortion, not remove it. A companion simulation (Brown & Heathcote, cited in preparation) reportedly finds the variation needed to produce significant bias "is often as little as one order of magnitude" in rate — a low bar, which is why real subject pools clear it.
This is the load-bearing distinction that separates Heathcote's specific claim from a loose "averaging hides individuals" slogan: it is a statement about the algebra of averaging exponentials with differing rates, not about aggregation in general. That specificity is exactly why the artifact does not transfer wholesale to the 2026 neural-scaling-law case — see claim-averaging-artifact-across-practice-and-scaling-laws-is-analogous-symptom-not-same-mechanism. It also sharpens the cross-domain bridge observation and gives teeth to the vault's standing hunch that firm-aggregated Wright's Law curves could carry the same bias.
Source
“Some variability among component learning rates is necessary for distortion of power and exponential function averages... The degree of distortion is proportional to the degree of learning rate variability. In particular, for averages over exponential functions, a power-like decrease in the relative learning rate will occur if the sample contains fast and slow learners.”
claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-26-do-heathcote-et-al-2000-and-neural-neural.md, 2026-08-01 · raw markdown