Whether a network forgets catastrophically or gracefully is governed by representational overlap, not by an intrinsic property of backpropagation
Given that catastrophic forgetting is a special regime rather than a universal fate (see claim-catastrophic-forgetting-is-a-structure-dependent-regime-not-universal), the literature offers a specific determinant of which regime appears: the degree to which the network's internal distributed representations overlap. French (1999), restating his own 1991–1992 work, holds that "catastrophic forgetting was largely a consequence of the overlap of internal distributed representations and that reducing this overlap would reduce catastrophic interference." The cause is located in the representation, not in gradient descent as such.
Two independent lines of evidence in the review support the mechanism. First, architecture: networks engineered toward semi-distributed or sparse (lower-overlap) internal codes reduce or eliminate catastrophic loss — yet the protection is not merely "less distributed is better." Convolution-correlation memory models (CHARM, TODAM) and Sparse Distributed Memory keep distributed representations and still avoid the cliff; their performance on previously learned information "declines gradually, rather than falling off abruptly, when learning new patterns." Second, task structure: French reports that pre-training a network on a regular, structured domain before sequential learning eliminates catastrophic interference in the later task (citing McRae & Hetherington 1993). Both point the same way — the graceful regime is bought by how the task and its representations are organized, not by abandoning distributed learning.
This is the mechanism underneath the feature-overlap intuition in claim-neural-nets-forget-along-human-like-power-law-curve (digit 8 keeps a partial trace because it shares strokes with 6 and 9) and the criterion by which claim-klines-mnist-drop-8-is-single-task-drift-not-a-disjoint-task-switch places Kline's protocol on the graceful side. It connects to the broader neuro-AI-parallel thread where representational geometry, not biology, does the explanatory work — cf. claim-critical-periods-arise-from-information-plasticity-not-biology.
Source
“catastrophic forgetting was largely a consequence of the overlap of internal distributed representations and that reducing this overlap would reduce catastrophic interference”
claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-13-is-catastrophic-vs-graceful-power-law-forgetting-a.md, 2026-07-18 · raw markdown