---
title: "Is GIFT's K-FAC gradient factor G = E[δδ⊤] the true Fisher (model's sampled predictive distribution) or the empirical Fisher (real training labels) — and does the difference weaken the GIFT–Amari 'same object' bridge?"
type: "question"
status: "answered"
date_raised: "2026-07-25T00:00:00.000Z"
answered_log: "2026-07-27 — [[claim-kunstner-empirical-fisher-not-equivalent-to-true-fisher]], [[claim-martens-grosse-kfac-defines-true-fisher-convention]], and [[claim-gift-g-factor-matches-empirical-fisher-not-true-fisher-convention]] resolve this, largely: direct reads of Kunstner et al. (2019), Martens & Grosse (2015), and GIFT's own method section show GIFT's G = E[δδ⊤] is built from real-label backpropagated gradients — structurally the empirical Fisher, not the true-Fisher convention Amari and canonical K-FAC specify. The 'same object' bridge softens to 'same K-FAC scheme, different Fisher variants.' Left open: the last claim is an inference from GIFT's silence on the sampling question, not a label GIFT applies to itself — a fully closed resolution would need an explicit author statement or a derivation of the conditions under which real-label G behaves like the true Fisher here."
tags: ["gradient-geometry","fisher-information","empirical-fisher","k-fac","natural-gradient","gift","primary-source-verification"]
---


Raised while promoting the capture behind
[[claim-gift-isotropy-transform-derived-from-fisher-kfac]] and
[[observation-gradient-geometry-shared-object-across-two-of-three]]. Those notes
establish that GIFT (arXiv:2607.07494) whitens gradients using the Fisher
information matrix in K-FAC-factored form — F_W ≈ A ⊗ G, with G = E[δδ⊤] on the
output-gradient side — the same formal object
[[claim-amari-1998-natural-gradient-fisher-steepest-descent|Amari's natural gradient]]
uses. But *which* Fisher is load-bearing for how tightly that identity holds.

Kunstner, Balles & Hennig (2019, arXiv:1905.12558, "Limitations of the Empirical
Fisher Approximation for Natural Gradient Descent") show the term "Fisher" is
used inconsistently across statistics and ML, and that the widely-used
**empirical Fisher** (gradient outer product from *actual training labels*) is
*not* in general the same matrix as the **true Fisher information** (from the
model's *own sampled predictive distribution*) that Amari's construction and the
original K-FAC convention (Martens & Grosse, 2015) assume. If GIFT's
G = E[δδ⊤] is computed from real labels, it is closer to the empirical Fisher,
and the "same object as Amari" claim is a near-twin rather than an identity.

**Why this is load-bearing, not a nice-to-have:** the observation note's central
verdict — GIFT and Amari share *one formal object* — rests on this being the
same Fisher, not two matrices sharing a name. It refines the strength of the
kept claim, so it belongs in the pile rather than as a note-body hedge.

**What to establish:**
1. Read GIFT's method section (§II–III) for whether the δ in G = E[δδ⊤] is
   sampled from the model's predictive distribution (true Fisher) or taken from
   real training targets (empirical Fisher). The capture's read did **not**
   specify which.
2. Cross-check against the K-FAC primary (Martens & Grosse, 2015), which both
   GIFT and the Kunstner critique build on, for the convention GIFT inherits.
3. If empirical: soften the identity in
   [[observation-gradient-geometry-shared-object-across-two-of-three]] and
   [[claim-gift-isotropy-transform-derived-from-fisher-kfac]] from "same object"
   to "same object up to the empirical-Fisher approximation."

Medium priority — the bridge stands either way (both are the Fisher/K-FAC
family, not GRADE's covariance-spectrum device), but the exact wording of "same
object" should not go evergreen until the factor is pinned.


## Progress log

- 2026-07-27 — [[claim-kunstner-empirical-fisher-not-equivalent-to-true-fisher]], [[claim-martens-grosse-kfac-defines-true-fisher-convention]], and [[claim-gift-g-factor-matches-empirical-fisher-not-true-fisher-convention]] resolve this, largely: direct reads of Kunstner et al. (2019), Martens & Grosse (2015), and GIFT's own method section show GIFT's G = E[δδ⊤] is built from real-label backpropagated gradients — structurally the empirical Fisher, not the true-Fisher convention Amari and canonical K-FAC specify. The 'same object' bridge softens to 'same K-FAC scheme, different Fisher variants.' Left open: the last claim is an inference from GIFT's silence on the sampling question, not a label GIFT applies to itself — a fully closed resolution would need an explicit author statement or a derivation of the conditions under which real-label G behaves like the true Fisher here.
