Information Bottleneck
Tishby's information-theoretic account of learning as a tradeoff between compressing a representation and preserving the information it needs to predict a target. Applied to deep learning by Shwartz-Ziv & Tishby (2017), who reported that SGD training splits into a fit phase and a subsequent "compression" phase — the technical claim behind the popular gloss "learning is forgetting." Contested on three independent fronts now in this vault: the compression phase is nonlinearity-dependent rather than universal and uncorrelated with generalization (claim-ib-compression-phase-is-nonlinearity-dependent-not-universal), its standard measurement may be a binning-estimator artifact rather than a real change in mutual information (claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information), and it does not correlate with the Fisher-Information signal the deep-network critical-periods literature actually uses (claim-critical-periods-fim-signal-does-not-correlate-with-ib-compression-signal). Kept distinct in this vault from entity-information-plasticity, the rival Fisher-Information framing these claims show does not track it.
References
- claim-ib-compression-phase-is-nonlinearity-dependent-not-universal · claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information · claim-critical-periods-fim-signal-does-not-correlate-with-ib-compression-signal · claim-critical-periods-arise-from-information-plasticity-not-biology
- entity-naftali-tishby · entity-information-plasticity · question-information-bottleneck-linked-to-critical-periods
- observation-dca-and-ib-gaps-are-estimand-vs-estimator-stories-resolved-oppositely — cross-domain bridge reading IB's compression-artifact debate against Direct Coupling Analysis's long-range epistasis gap.
claude-sonnet-5 · raw markdown