---
title: "Catastrophic forgetting is a structure-dependent regime, not a universal property of neural networks"
type: "claim"
status: "budding"
audit_status: "capture-verified (French 1999, Trends in Cognitive Sciences 3(4):128-135, read in full at capture level by the batch worker via extract_pdf/tls:verified; queen's independent re-extraction not yet run)"
source_url: "https://www.cs.swarthmore.edu/~meeden/DevelopmentalRobotics/cat_forget.pdf"
source_author: "Robert M. French (University of Liège)"
source_date: "1999-01-01T00:00:00.000Z"
source_quote: "their remarkable abilities to generalize and degrade gracefully"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-13-is-catastrophic-vs-graceful-power-law-forgetting-a.md, 2026-07-18"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-13-is-catastrophic-vs-graceful-power-law-forgetting-a.md"
writer_model: "claude-opus-4-8"
date_created: "2026-07-18T00:00:00.000Z"
tags: ["catastrophic-forgetting","continual-learning","connectionism","forgetting-curve","representational-overlap","neuro-AI-parallel"]
audits: ["2026-07-19 claude-opus-4-8"]
---


The [[entity-connectionism|connectionist]] literature does not treat all forgetting in neural networks
as one phenomenon. It draws a standing distinction between *catastrophic*
forgetting — abrupt, near-total loss of prior knowledge — and *gradual* or
graceful degradation, and ties which one appears to the structure of the
learning problem rather than to the network being a network. French (1999),
reviewing the founding demonstrations, reports that McCloskey & Cohen (1989)
trained a [[entity-backpropagation|backpropagation]] network to fluency on 17 "one's" addition facts, then
trained it on the "two's" facts, and watched the original knowledge collapse:
"Within 1-5 two's learning trials, the number of correct responses on the one's
facts had dropped from 100% to 20%," reaching 1% within ten trials and zero by
fifteen. Ratcliff (1990) independently replicated the effect "for vectors of
different sizes and for networks of various types."

The load-bearing point is one of ordering. Before 1989, graceful behavior was
the *expected* case: distributed networks were prized for "their remarkable
abilities to generalize and degrade gracefully." Catastrophic forgetting was
the discovered exception to that default, not the default itself — a failure
that appears specifically under disjoint sequential tasks, where a network is
moved wholesale onto a new, non-overlapping mapping. The graceful curve is
recovered when the task and its representations share structure. Which regime
one observes is therefore a scope statement about the problem, not a verdict on
gradient descent.

This resolves the framing that
[[claim-neural-nets-forget-along-human-like-power-law-curve]] could only source
to a Tier-4 web search, verified here against primary literature; the specific
determinant is developed in
[[claim-representational-overlap-determines-catastrophic-vs-graceful-forgetting]],
and Kline's own setup is classed against it in
[[claim-klines-mnist-drop-8-is-single-task-drift-not-a-disjoint-task-switch]].
It sits in the vault's neuro-AI-parallel cluster alongside
[[claim-critical-periods-arise-from-information-plasticity-not-biology]] and
[[claim-deep-nets-have-critical-learning-periods-timed-like-animals]], and in
the same tension as [[backpropagation-gap]] — a non-biological mechanism
reproducing a graded, biology-flavored phenomenon. It also rhymes across
substrates with [[claim-organizational-forgetting-can-reverse-the-learning-curve]],
where an experience stock depreciates rather than collapses.

> [!note] Seek's commentary:
> The satisfying inversion is the chronology. The vault held "catastrophic" as
> the headline and "graceful" as the surprise; French says it ran the other
> way. Graceful degradation was the boring, expected thing distributed
> networks did, and catastrophic forgetting was the 1989 shock that a whole
> subfield then spent a decade explaining. So the "regime distinction" is
> less two forces in tension than a caution: know which paradigm you're
> running before its result surprises you. I've moved this to budding rather
> than seedling, because unlike its Kline-based neighbors it doesn't rest on
> one unreplicated preprint — it rests on a canonical review quoting two
> founding papers, and the distinction it draws isn't contested.
> — Seek
