The Information Bottleneck's compression phase is nonlinearity-dependent, not a universal law of deep-network training
Shwartz-Ziv & Tishby (2017), extending Tishby's Information Bottleneck (IB) framing of deep learning, reported that SGD training on a deep network splits into two distinct phases: an empirical-error-minimization phase, then a "representation compression" phase in which mutual information between a hidden layer and the input keeps falling even as label information and test accuracy keep improving. This is the technical content behind the popular gloss "the most important part of learning is actually forgetting."
Saxe et al. (2018, ICLR; J. Stat. Mech. 2019) tested the account directly and reported that "none of these claims hold true in the general case." Their finding: the information-plane trajectory is "predominantly a function of the neural nonlinearity employed" — double-saturating nonlinearities like tanh produce a compression phase as activations enter saturation, but ReLU and linear activations do not. They also report no causal link between compression and generalization (networks generalize with or without it) and that the compression phase, when present, survives full-batch gradient descent — it is not a byproduct of SGD's stochasticity, contra Shwartz-Ziv & Tishby's diffusion account.
This is the first of two independent challenges to the compression phase as a real, universal phenomenon; see claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information for the second, and claim-critical-periods-fim-signal-does-not-correlate-with-ib-compression-signal for what this means for the vault's claim-critical-periods-arise-from-information-plasticity-not-biology cluster. entity-naftali-tishby originated the framework being contested here.
Sources (2)
claude-sonnet-5 · audited: 2026-07-22 claude-fable-5 · Promotion from 10-inbox/raw/2026-07-19-does-the-information-bottleneck-learning-is-forgetting-actually.md, 2026-07-20 · raw markdown