---
id: "20260719-0202-does-the-information-bottleneck"
title: "Does the Information Bottleneck ('learning is forgetting') actually explain deep-network critical periods, and is its compression phase even universal?"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-ib-compression-phase-is-nonlinearity-dependent-not-universal.md","30-notes/claim-ib-compression-may-be-a-binning-artifact-not-real-mutual-information.md","30-notes/claim-critical-periods-fim-signal-does-not-correlate-with-ib-compression-signal.md","40-entities/entity-naftali-tishby.md","40-entities/entity-alessandro-achille.md","40-entities/entity-stefano-soatto.md","40-entities/entity-information-bottleneck.md","40-entities/entity-information-plasticity.md"]
not_promoted: ["Achille & Soatto (2018, JMLR) as an independent claim — flagged in Further leads as the FIM-to-activation-information bridge paper, but not independently read this session; a lead for a future capture, not a sourced claim to promote now.","Tishby, Pereira & Bialek (1999) as an independent claim — the foundational IB-method paper; not read this session, no primary quote to ground a note on.","Tishby & Zaslavsky (2015) as an independent claim — represented here only via Shwartz-Ziv & Tishby (2017)'s restatement; not independently re-read, so folded as background rather than promoted on its own provenance.","Quanta Magazine's 'learning is forgetting' framing quote — Tier 2, used only as color inside claim-ib-compression-phase-is-nonlinearity-dependent-not-universal; not promoted as its own claim since it is a popularization of Claim 1's technical content, not a distinct assertion.","Whether any post-2019 paper derives critical periods directly from an IB objective — not searched this session; a candidate future capture, not a claim, so not routed to 50-questions/ (no kept claim rests on this doubt; it is a nice-to-have completeness check, not load-bearing).","Concept entities Information Plane, compression phase (as a named sub-concept), and Fisher Information Matrix — flagged as entity candidates in the capture but folded into the connects_to fields of the Information Bottleneck / Information Plasticity hubs rather than given standalone entity pages, to avoid stub-page flood over three closely related terms.","Person entities Ravid Shwartz-Ziv, Andrew Saxe, and Ziv Goldfeld — each real and load-bearing for the one paper they lead in this dispute, but each appears only once in the vault so far with no independent research profile established beyond that single paper (Saxe's other cited work, loss landscapes, is not otherwise represented in the vault); left as plain-text mentions in the claim-notes rather than promoted to entity hubs. Recoverable later if any of the three recurs."]
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-07-19T00:00:00.000Z"
provenance: "batch run, 2026-07-19; direct answer to 50-questions/question-information-bottleneck-linked-to-critical-periods.md"
derived_from: []
tags: ["information-bottleneck","critical-periods","tishby","saxe","information-plasticity","fisher-information","compression-phase","deep-learning","verification"]
sources: [{"source_url":"https://arxiv.org/pdf/1703.00810","source_author":"Ravid Shwartz-Ziv, Naftali Tishby","source_date":"2017-03 (v3 2017-04-29)","source_tier":1},{"source_url":"https://artemyk.github.io/assets/pdf/papers/Saxe%20et%20al_2018_ICLR.pdf","source_author":"Andrew M. Saxe, Yamini Bansal, Joel Dapello, Madhu Advani, Artemy Kolchinsky, Brendan D. Tracey, David D. Cox","source_date":"2018 (ICLR 2018; journal version J. Stat. Mech. 2019-12-20)","source_tier":1},{"source_url":"https://arxiv.org/pdf/1810.05728","source_author":"Ziv Goldfeld, Ewout van den Berg, Kristjan Greenewald, Igor Melnyk, Nam Nguyen, Brian Kingsbury, Yury Polyanskiy","source_date":"2019 (ICML 2019)","source_tier":1},{"source_url":"https://arxiv.org/abs/1711.08856","source_author":"Alessandro Achille, Matteo Rovere, Stefano Soatto","source_date":"2019-02-25T00:00:00.000Z","source_tier":1},{"source_url":"https://www.quantamagazine.org/new-theory-cracks-open-the-black-box-of-deep-learning-20170921/","source_author":"Natalie Wolchover (quoting Naftali Tishby)","source_date":"2017-09-21T00:00:00.000Z","source_tier":2}]
---


*This capture directly answers [[question-information-bottleneck-linked-to-critical-periods]], the open question routed from [[claim-critical-periods-arise-from-information-plasticity-not-biology]]. Three core claims, tight to the question. TLS verified on all PDF fetches (extract_pdf); no elevated-suspicion sources encountered. No safety-flag-triggering content found on any fetched page.*

---

## Claim 1: The Information Bottleneck theory's own claims about a universal fit-then-compress cycle were empirically refuted by a follow-up study — compression is nonlinearity-dependent, not causally tied to generalization, and not caused by SGD stochasticity

**Claim type:** specific technical-mechanism claim → Tier 1–2 required. Met: Tier 1, both sides of the dispute read as primaries.

Tishby & Zaslavsky (2015) proposed analyzing deep networks in an "Information Plane" — tracking mutual information between each layer and the input/output — and framed training as an Information Bottleneck (IB) tradeoff. Shwartz-Ziv & Tishby (2017) then reported an empirical two-phase signature for this process: "Our analysis reveals, for the first time to our knowledge, that the Stochastic Gradient De[s]cent (SGD) optimization, commonly used in Deep Learning, has two different and distinct phases: empirical error minimization (ERM) and representation compression... In the compression phase, the fluctuations of the gradients are much larger than their means, and the weights change essentially as Wiener processes, or random diffusion." This compression phase — training epochs spent shrinking mutual information between a hidden layer and the input, even as label information and test performance keep improving — is the technical content behind the popular gloss "the most important part of learning is actually forgetting" (see Claim framing note below).

Saxe et al. (2018, ICLR; published J. Stat. Mech. 2019) tested this account directly and reported the opposite of a universal law. Their abstract states the IB theory of deep learning "makes three specific claims: first, that deep networks undergo two distinct phases consisting of an initial fitting phase and a subsequent compression phase; second, that the compression phase is causally related to the excellent generalization performance of deep networks; and third, that the compression phase occurs due to the diffusion-like behavior of stochastic gradient descent. Here we show that none of these claims hold true in the general case." Instead: "the information plane trajectory is predominantly a function of the neural nonlinearity employed: double-sided saturating nonlinearities like tanh yield a compression phase as neural activations enter the saturation regime, but linear activation functions and single-sided saturating nonlinearities like the widely used ReLU in fact do not." They also report "no evident causal connection between compression and generalization: networks that do not compress are still capable of generalization, and vice versa," and that the compression phase, when present, replicates under full-batch gradient descent — i.e. it "does not arise from stochasticity in training."

**Source:**
- source_url: https://arxiv.org/pdf/1703.00810 (Shwartz-Ziv & Tishby 2017, extract_pdf, tls: verified)
- grounding quote: *"the Stochastic Gradient De[s]cent (SGD) optimization... has two different and distinct phases: empirical error minimization (ERM) and representation compression"*
- source_url: "https://artemyk.github.io/assets/pdf/papers/Saxe%20et%20al_2018_ICLR.pdf" (Saxe et al. 2018, extract_pdf, tls: verified; author-hosted copy of the ICLR 2018 conference paper, later published in J. Stat. Mech. 2019(12):124020)
- grounding quote: *"Here we show that none of these claims hold true in the general case... the information plane trajectory is predominantly a function of the neural nonlinearity employed: double-sided saturating nonlinearities like tanh yield a compression phase... but linear activation functions and single-sided saturating nonlinearities like the widely used ReLU in fact do not."*
- source_tier: 1 (both)

---

## Claim 2: A third primary source shows the measured "compression" itself may be a binning/estimation artifact reflecting geometric clustering, not a genuine change in mutual information — deepening the doubt about whether a compression phase is even a coherent, real phenomenon

**Claim type:** specific technical-mechanism claim → Tier 1–2 required. Met: Tier 1.

Goldfeld et al. (2019, ICML) revisit the mutual-information estimator both Shwartz-Ziv & Tishby (2017) and Saxe et al. (2018) used — a per-neuron discretization ("binning") of continuous activations — and show it is mathematically degenerate for the deterministic networks being studied: "in deterministic DNNs with strictly monotone nonlinearities (e.g., tanh or sigmoid) the true mutual information I(X; T_ℓ) is provably either infinite (continuous X) or a constant (discrete X). Therefore, the fluctuations of I(X; Bin(T_ℓ)) observed during DNN training by (Shwartz-Ziv & Tishby, 2017; Saxe et al., 2018) must be due to estimation errors rather than changes in mutual information." Their own analysis of what the binned proxy is actually tracking: "compression, i.e. reduction in I(X; T_ℓ) over the course of training, is driven by progressive geometric clustering of the representations of samples from the same class." They report this clustering occurs "in purely deterministic DNNs, where I(X; T_ℓ) is provably vacuous" too — i.e. even where "true" mutual information cannot change at all, the clustering signature that had been read as "compression" still appears, and the paper states this "provide[s] new evidence that compression and generalization may not be causally related."

This bears on the second half of the topic question directly: not only is the *causal* role of compression in generalization contested (Claim 1), but whether the standard measurement method was tracking a real information-theoretic quantity, as opposed to a geometric side effect with no principled compression interpretation, is itself in question.

**Source:**
- source_url: https://arxiv.org/pdf/1810.05728 (Goldfeld et al. 2019, extract_pdf, tls: verified; ICML 2019 proceedings paper)
- grounding quote: *"in deterministic DNNs with strictly monotone nonlinearities... the true mutual information I(X; T_ℓ) is provably either infinite (continuous X) or a constant (discrete X). Therefore, the fluctuations of I(X; Bin(T_ℓ)) observed during DNN training by (Shwartz-Ziv & Tishby, 2017; Saxe et al., 2018) must be due to estimation errors rather than changes in mutual information."*
- source_tier: 1

---

## Claim 3: The founding deep-network critical-period paper grounds its result in a different information quantity than Tishby's IB compression signal, and explicitly reports its own metric does not correlate with that signal — the IB link is shared vocabulary, not a demonstrated mechanism

**Claim type:** specific technical-mechanism claim → Tier 1–2 required. Met: Tier 1, primary paper, direct re-extraction.

Achille, Rovere & Soatto (2019) — the paper that established critical periods in deep networks, already covered in [[claim-critical-periods-arise-from-information-plasticity-not-biology]] and [[claim-founding-dnn-critical-period-paper-validated-on-animal-data-only]] — build their account on the Fisher Information Matrix (FIM) of the network's *weights*, not on Shannon mutual information of the *activations* that Tishby's IB framework tracks. The authors state this distinction themselves: "The existence of two distinct phases of training has been observed and discussed by Shwartz-Ziv & Tishby (2017), although their analysis builds on the (Shannon) information of the activations, rather than the (Fisher) information in the weights." They go further and report that the specific statistic Shwartz-Ziv & Tishby use to mark the ERM→compression transition — gradient covariance — does not track their own critical-period sensitivity signal: "it must be noted that the FIM is computed using the gradients with respect to the model prediction, not to the ground truth label, leading to important qualitative differences. In Figure 6, we show that the covariance and norm of the gradients exhibit no clear trends during training with and without deficits, and, therefore, unlike the FIM, do not correlate with the sensitivity to critical periods."

The paper does draw an indirect bridge — citing Achille & Soatto (2018, JMLR), "a connection between our FIM analysis and the information in the activations can be established based on [that] work, which shows that the FIM of the weights can be used to bound the information in the activations" — but this is an inferential link through a separate paper's bound, not a demonstration that the critical-period phenomenon is produced by, or requires, an IB-style compression phase. The paper's own vocabulary choices ("forgetting," "bottleneck crossings" for a loss-landscape curvature bottleneck, distinct from Tishby's information bottleneck) read as an explicit borrowing of framing rather than a claim of shared mechanism — consistent with the same paper's disclaimer, recorded separately, that it does not claim DNNs are "necessarily a valid model of neurobiological information processing" ([[claim-achille-soatto-disclaim-dnn-as-valid-model-of-biology]]).

**Central question status:** `[unverified in the strong sense — could not confirm the Information Bottleneck compression phase as the mechanism behind deep-network critical periods; the founding paper's own analysis dissociates its critical-period metric from Tishby's compression signal]`. Combined with Claims 1–2, the honest state of the literature as searched is: the IB compression phase is not established as universal (Saxe et al.), its very measurement is disputed as a possible artifact (Goldfeld et al.), and the paper that would need to derive critical periods *from* IB instead uses a different quantity that its own authors show does not correlate with the IB-style compression signature. The "IB explains critical periods" framing is not confirmed by primary sources; it survives only as a resemblance in narrative shape (rise-then-fall, "forgetting") between two information-theoretic quantities shown not to move together.

**Source:**
- source_url: https://arxiv.org/abs/1711.08856 (PDF re-extracted via extract_pdf, tls: verified)
- source_author: Alessandro Achille, Matteo Rovere, Stefano Soatto
- source_date: 2019-02-25
- source_tier: 1
- grounding quote 1: *"The existence of two distinct phases of training has been observed and discussed by Shwartz-Ziv & Tishby (2017), although their analysis builds on the (Shannon) information of the activations, rather than the (Fisher) information in the weights."*
- grounding quote 2: *"the covariance and norm of the gradients exhibit no clear trends during training with and without deficits, and, therefore, unlike the FIM, do not correlate with the sensitivity to critical periods."*

---

## Further leads

- **Achille & Soatto (2018)**, "Emergence of Invariance and Disentanglement in Deep Representations," *JMLR* 19(1):1947–1980 — the paper Achille, Rovere & Soatto (2019) cite as the bridge between FIM-of-weights and information-in-activations; not independently read this session, would need its own check to see whether it derives critical periods from an IB-style objective more directly than the 2019 paper does.
- **Tishby, Pereira & Bialek (1999)**, the original Information Bottleneck method paper — foundational IB reference cited by both Tishby & Zaslavsky (2015) and Shwartz-Ziv & Tishby (2017); not read this session, would ground the "IB bound" concept itself rather than its deep-learning application.
- **Tishby & Zaslavsky (2015)**, "Deep Learning and the Information Bottleneck Principle," arXiv:1503.02406 — the earlier paper that first proposed studying DNNs in the Information Plane; located but not independently re-read this session (its claims are represented here via Shwartz-Ziv & Tishby 2017's restatement of them).
- Quanta Magazine's framing that "the most important part of learning is actually forgetting" is a direct quote from Tishby in a 2017 interview (Natalie Wolchover, byline confirmed, no addressed-to-AI or override language found on the page) but is a popularization, not the technical claim itself — the technical claim it summarizes is the ERM/compression two-phase account in Claim 1, which Saxe et al. (2018) and Goldfeld et al. (2019) subsequently contest. Recorded here as `source_tier: 2` (direct quote from a named source via a journalism venue) and used only for the framing/definitional note, not for any quantitative or mechanism claim.
- Whether any paper *after* 2019 has attempted to derive critical periods directly from an IB objective (rather than from FIM) was not searched this session — a natural next hop if the vault wants a fully exhaustive answer rather than a founding-paper-scoped one.

## Entity candidates

- Naftali Tishby — person — originator of the Information Bottleneck principle and its deep-learning application; central figure across the whole cluster, dead in 2021, worth a bio/context page.
- Ravid Shwartz-Ziv — person — co-author of the 2017 "Opening the Black Box" paper that produced the empirical two-phase/compression claim being contested.
- Andrew Saxe — person — lead author of the 2018 refutation paper; also known for other deep-learning theory work (loss landscapes).
- Alessandro Achille — person — co-author of the founding DNN critical-periods paper and the Information Plasticity framework; recurring figure in this vault cluster.
- Stefano Soatto — person — co-author on both the critical-periods paper and the Information Plasticity/invariance work; also already flagged in [[claim-achille-soatto-disclaim-dnn-as-valid-model-of-biology]].
- Ziv Goldfeld — person — lead author of the 2019 ICML paper resolving the compression-phase measurement dispute via geometric clustering; not yet represented in the vault.
- Information Bottleneck (principle) — concept — the general information-theoretic training/compression framework this whole cluster revolves around; currently no dedicated vault note despite being load-bearing for several claim-notes.
- Information Plasticity — concept — Achille/Rovere/Soatto's Fisher-Information-based rival framing to IB; already used in existing claim titles but may warrant its own concept page given Claim 3's finding that it does not correlate with the IB compression metric.
- Fisher Information Matrix (in deep learning) — concept — the specific technical quantity (of weights, not activations) underlying Information Plasticity; distinguishing it clearly from Shannon mutual information may be useful as its own note given how often the two get conflated in secondary sources.
- Information Plane — concept/term — Tishby & Zaslavsky's visualization method (mutual-information-with-input vs. mutual-information-with-output per layer); the shared analytical surface all three primary papers argue over.
- Compression phase (IB) — concept — the specific contested phenomenon; worth its own concept note given how much dispute concentrates on whether it's (a) universal, (b) causally linked to generalization, (c) a genuine information-theoretic quantity at all.
