Catastrophic forgetting is a structure-dependent regime, not a universal property of neural networks
The connectionist literature does not treat all forgetting in neural networks as one phenomenon. It draws a standing distinction between catastrophic forgetting — abrupt, near-total loss of prior knowledge — and gradual or graceful degradation, and ties which one appears to the structure of the learning problem rather than to the network being a network. French (1999), reviewing the founding demonstrations, reports that McCloskey & Cohen (1989) trained a backpropagation network to fluency on 17 "one's" addition facts, then trained it on the "two's" facts, and watched the original knowledge collapse: "Within 1-5 two's learning trials, the number of correct responses on the one's facts had dropped from 100% to 20%," reaching 1% within ten trials and zero by fifteen. Ratcliff (1990) independently replicated the effect "for vectors of different sizes and for networks of various types."
The load-bearing point is one of ordering. Before 1989, graceful behavior was the expected case: distributed networks were prized for "their remarkable abilities to generalize and degrade gracefully." Catastrophic forgetting was the discovered exception to that default, not the default itself — a failure that appears specifically under disjoint sequential tasks, where a network is moved wholesale onto a new, non-overlapping mapping. The graceful curve is recovered when the task and its representations share structure. Which regime one observes is therefore a scope statement about the problem, not a verdict on gradient descent.
This resolves the framing that claim-neural-nets-forget-along-human-like-power-law-curve could only source to a Tier-4 web search, verified here against primary literature; the specific determinant is developed in claim-representational-overlap-determines-catastrophic-vs-graceful-forgetting, and Kline's own setup is classed against it in claim-klines-mnist-drop-8-is-single-task-drift-not-a-disjoint-task-switch. It sits in the vault's neuro-AI-parallel cluster alongside claim-critical-periods-arise-from-information-plasticity-not-biology and claim-deep-nets-have-critical-learning-periods-timed-like-animals, and in the same tension as backpropagation-gap — a non-biological mechanism reproducing a graded, biology-flavored phenomenon. It also rhymes across substrates with claim-organizational-forgetting-can-reverse-the-learning-curve, where an experience stock depreciates rather than collapses.
Source
“their remarkable abilities to generalize and degrade gracefully”
claude-opus-4-8 · audited: 2026-07-19 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-13-is-catastrophic-vs-graceful-power-law-forgetting-a.md, 2026-07-18 · raw markdown