Is GIFT's K-FAC gradient factor G = E[δδ⊤] the true Fisher (model's sampled predictive distribution) or the empirical Fisher (real training labels) — and does the difference weaken the GIFT–Amari 'same object' bridge?
Raised while promoting the capture behind claim-gift-isotropy-transform-derived-from-fisher-kfac and observation-gradient-geometry-shared-object-across-two-of-three. Those notes establish that GIFT (arXiv:2607.07494) whitens gradients using the Fisher information matrix in K-FAC-factored form — F_W ≈ A ⊗ G, with G = E[δδ⊤] on the output-gradient side — the same formal object Amari's natural gradient uses. But which Fisher is load-bearing for how tightly that identity holds.
Kunstner, Balles & Hennig (2019, arXiv:1905.12558, "Limitations of the Empirical Fisher Approximation for Natural Gradient Descent") show the term "Fisher" is used inconsistently across statistics and ML, and that the widely-used empirical Fisher (gradient outer product from actual training labels) is not in general the same matrix as the true Fisher information (from the model's own sampled predictive distribution) that Amari's construction and the original K-FAC convention (Martens & Grosse, 2015) assume. If GIFT's G = E[δδ⊤] is computed from real labels, it is closer to the empirical Fisher, and the "same object as Amari" claim is a near-twin rather than an identity.
Why this is load-bearing, not a nice-to-have: the observation note's central verdict — GIFT and Amari share one formal object — rests on this being the same Fisher, not two matrices sharing a name. It refines the strength of the kept claim, so it belongs in the pile rather than as a note-body hedge.
What to establish:
- Read GIFT's method section (§II–III) for whether the δ in G = E[δδ⊤] is sampled from the model's predictive distribution (true Fisher) or taken from real training targets (empirical Fisher). The capture's read did not specify which.
- Cross-check against the K-FAC primary (Martens & Grosse, 2015), which both GIFT and the Kunstner critique build on, for the convention GIFT inherits.
- If empirical: soften the identity in observation-gradient-geometry-shared-object-across-two-of-three and claim-gift-isotropy-transform-derived-from-fisher-kfac from "same object" to "same object up to the empirical-Fisher approximation."
Medium priority — the bridge stands either way (both are the Fisher/K-FAC family, not GRADE's covariance-spectrum device), but the exact wording of "same object" should not go evergreen until the factor is pinned.
Progress log
- 2026-07-27 — claim-kunstner-empirical-fisher-not-equivalent-to-true-fisher, claim-martens-grosse-kfac-defines-true-fisher-convention, and claim-gift-g-factor-matches-empirical-fisher-not-true-fisher-convention resolve this, largely: direct reads of Kunstner et al. (2019), Martens & Grosse (2015), and GIFT's own method section show GIFT's G = E[δδ⊤] is built from real-label backpropagated gradients — structurally the empirical Fisher, not the true-Fisher convention Amari and canonical K-FAC specify. The 'same object' bridge softens to 'same K-FAC scheme, different Fisher variants.' Left open: the last claim is an inference from GIFT's silence on the sampling question, not a label GIFT applies to itself — a fully closed resolution would need an explicit author statement or a derivation of the conditions under which real-label G behaves like the true Fisher here.