Kunstner, Balles & Hennig (2019) formally distinguish the true Fisher from the empirical Fisher and show the empirical Fisher is not, in general, a valid estimate of it
Kunstner, Balles & Hennig define the Fisher information matrix as an expectation over the model's own predictive distribution: "F(θ) := Σₙ E_{p_θ(y|xₙ)}[∇_θ log p_θ(y|xₙ) ∇_θ log p_θ(y|xₙ)ᵀ]." They separately name the widely-used substitute — built from real training labels instead — the "empirical Fisher": "F̃(θ) := Σₙ ∇_θ log p_θ(yₙ|xₙ) ∇_θ log p_θ(yₙ|xₙ)ᵀ," where yₙ is the actual label. The paper is explicit that these are not interchangeable by construction: yₙ is a training label, not a sample from p_θ(y|xₙ), so "the empirical Fisher is not an empirical (i.e. Monte Carlo) estimate of the Fisher," despite the name. Convergence between the two "depends on how close the model p_θ(y|xₙ) is to the true data-generating distribution p(y|xₙ)" — they coincide only near a well-fit optimum, not generally during training — and the paper's central argument is that the conditions for near-equivalence "are unlikely to be met in practice."
This is the formal vocabulary the vault's gradient-geometry cluster now leans on to check whether a given K-FAC-style Fisher approximation (K-FAC) is using the true or empirical variant: claim-martens-grosse-kfac-defines-true-fisher-convention and claim-gift-g-factor-matches-empirical-fisher-not-true-fisher-convention.
Source
“yₙ is a training label and not a sample from the model's predictive distribution p_θ(y|xₙ). Therefore, and contrary to what its name suggests, the empirical Fisher is not an empirical (i.e. Monte Carlo) estimate of the Fisher.”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-07-26-is-gifts-k-fac-gradient-factor-g-eδδ.md, 2026-07-27 · raw markdown