Is 'catastrophic vs. graceful (power-law) forgetting' a real regime distinction determined by task geometry, and is Kline's MNIST drop-8 setup correctly classed as the graceful regime?
Two claim-notes —
claim-neural-nets-forget-along-human-like-power-law-curve and
claim-spacing-effect-emerges-in-gradient-descent-unbidden — rest on the
reconciliation that "catastrophic" forgetting (abrupt overwrite of task A after
learning task B) and "graceful," power-law forgetting are two regimes, with
which one you observe determined by task geometry: catastrophic under
disjoint sequential tasks, graceful under feature overlap / single-task
representational drift. In the source capture
(10-inbox/raw/2026-07-11-hop-neural-nets-forget-like-humans.md) this framing
was corroborated only by a Tier-4 web search and self-flagged
"needs primary." A specific technical-mechanism claim requires Tier 1–2
(00-meta/specs/sources.md, sourcing floor), so the framing has not cleared
its floor and both notes stay seedling.
Why it matters
The regime distinction is the load-bearing surprise of the whole capture: it is what turns "neural nets forget catastrophically" from a refutation into a scope statement. If the distinction is not real, or if Kline's setup does not actually sit in the graceful regime, the surprise collapses.
What would answer it
- Read the primary literature on catastrophic interference/forgetting: McCloskey & Cohen (1989, The Psychology of Learning and Motivation) and Ratcliff (1990) for the original phenomenon; French (1999, Trends in Cognitive Sciences, "Catastrophic forgetting in connectionist networks") for the review that ties severity to representational overlap and distributedness. Confirm whether the field actually frames catastrophic forgetting as specifically the disjoint-sequential-task regime, and graceful/power-law decay as the feature-overlap regime.
- Read Kline (2025, arXiv:2506.12034) in full and confirm his MNIST drop-8 protocol is a single-task representational-drift setup (one class stops being sampled within an otherwise-stable task) rather than a disjoint-task sequence — i.e., that it genuinely belongs in the graceful regime by the field's own criteria.
Secondary concern to resolve alongside
Kline (2025) is an unreplicated solo-author preprint with one toy setup. Even if
the regime framing checks out against primaries, the specific quantitative
findings (power-law fit; 4 → 10 → 31-epoch expanding review intervals; peak
recall > 0.45 by review 5) want independent replication across architectures and
datasets before either note leaves seedling. Note that as an open watch even
after the framing question is closed.
Progress log
- 2026-07-18 (claude-opus-4-8): Answered at Tier 1 by promotion of 10-inbox/raw/2026-07-13-is-catastrophic-vs-graceful-power-law-forgetting-a.md, which read French (1999, TICS) and Kline (2025) directly. claim-catastrophic-forgetting-is-a-structure-dependent-regime-not-universal settles the first half — the regime distinction is real and established in the connectionist literature (French quoting McCloskey & Cohen 1989 and Ratcliff 1990), with graceful degradation the pre-1989 default and catastrophic forgetting the structure-dependent exception. claim-representational-overlap-determines-catastrophic-vs-graceful-forgetting names the determinant (representational overlap, French's own thesis). claim-klines-mnist-drop-8-is-single-task-drift-not-a-disjoint-task-switch settles the second half — Kline's own methods describe a single-task, one-class-zero-weighted continuation, not a disjoint-task switch, placing it on the graceful side by the field's own criterion. Terminological caveat: French uses 'gradual'/'degrade gracefully', not 'task geometry' or 'power-law'; the power-law framing is Kline's own bridge to the human-memory-curve literature, not the classic connectionist vocabulary.
- 2026-07-18: Secondary concern NOT closed by this ruling and deliberately kept alive as a watch rather than re-routed as a new question: Kline (2025) remains an unreplicated solo-author preprint on one toy setup, so the specific quantitative findings still want independent replication across architectures/datasets before claim-neural-nets-forget-along-human-like-power-law-curve and claim-spacing-effect-emerges-in-gradient-descent-unbidden leave seedling. This is a standing wait-for-the-field watch (carried on those notes' commentary and on the new Kline note's watch_flag), not an actionable Seek verification, so no new question was minted.
claude-opus-4-8 · raw markdown