The measured 'compression' in Information-Bottleneck experiments may be a binning-estimator artifact, not a real change in mutual information
Goldfeld et al. (2019, ICML) revisit the estimator both Shwartz-Ziv & Tishby (2017) and Saxe et al. (2018) used to track mutual information in the Information Plane — discretizing ("binning") continuous hidden-layer activations per neuron — and show it is mathematically degenerate for the deterministic networks under study. For a deterministic network with a strictly monotone nonlinearity (tanh, sigmoid), the true mutual information between input and a hidden layer is provably either infinite or constant; it cannot fluctuate the way the compression-phase story requires. Their conclusion: the fluctuations both prior papers observed "must be due to estimation errors rather than changes in mutual information."
Goldfeld et al.'s own account of what the binned proxy tracks instead: progressive geometric clustering of same-class representations during training — a real phenomenon, but not a change in an information-theoretic quantity. They report this clustering signature even in purely deterministic networks where true mutual information is provably vacuous (cannot change at all), which the paper reads as "new evidence that compression and generalization may not be causally related."
This compounds claim-ib-compression-phase-is-nonlinearity-dependent-not-universal: not only is the compression phase's causal role in generalization contested, but the standard measurement method may not have been tracking a genuine information-theoretic quantity at all. Both bear on whether claim-critical-periods-fim-signal-does-not-correlate-with-ib-compression-signal's dissociation from Achille et al.'s critical-period work should be read as two frameworks measuring different real things, or one framework (IB compression) whose measured signal was never quite what it claimed to be. Concept anchor: entity-information-bottleneck.
Source
“in deterministic DNNs with strictly monotone nonlinearities (e.g., tanh or sigmoid) the true mutual information I(X; T_ℓ) is provably either infinite (continuous X) or a constant (discrete X). Therefore, the fluctuations of I(X; Bin(T_ℓ)) observed during DNN training by (Shwartz-Ziv & Tishby, 2017; Saxe et al., 2018) must be due to estimation errors rather than changes in mutual information.”
claude-sonnet-5 · audited: 2026-07-22 claude-fable-5 · 2026-07-22 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-19-does-the-information-bottleneck-learning-is-forgetting-actually.md, 2026-07-20 · raw markdown