---
title: "Kunstner, Balles & Hennig (2019) formally distinguish the true Fisher from the empirical Fisher and show the empirical Fisher is not, in general, a valid estimate of it"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://arxiv.org/pdf/1905.12558"
source_title: "Limitations of the Empirical Fisher Approximation for Natural Gradient Descent"
source_author: "Frederik Kunstner, Lukas Balles, Philipp Hennig"
source_date: 2019
source_venue: "arXiv:1905.12558 (NeurIPS 32, 2019, 'Limitations of the Empirical Fisher Approximation for Natural Gradient Descent'; v3 revised 2020-06-08)"
source_quote: "yₙ is a training label and not a sample from the model's predictive distribution p_θ(y|xₙ). Therefore, and contrary to what its name suggests, the empirical Fisher is not an empirical (i.e. Monte Carlo) estimate of the Fisher."
source_tier: 1
audit_status: "capture-verified (read directly via extract_pdf, 2026-07-26 session, tls: verified; not independently re-fetched by a second model this promotion — headless run, no network access. Held at seedling.) — 2026-07-28 cross-model audit (auditor claude-fable-5; writer claude-sonnet-5): independently re-fetched arXiv:1905.12558 v3 via extract_pdf (tls verified). All three quoted passages verified verbatim (pp. 2 and §3.2); F(θ)/F̃(θ) definitions match Eqs. 2–3/6–7. Clean — no corrections."
provenance: "Promotion from 10-inbox/raw/2026-07-26-is-gifts-k-fac-gradient-factor-g-eδδ.md, 2026-07-27"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-26-is-gifts-k-fac-gradient-factor-g-eδδ.md"
date_created: "2026-07-27T00:00:00.000Z"
tags: ["gradient-geometry","fisher-information","empirical-fisher","natural-gradient","primary-source-verification"]
---


Kunstner, Balles & Hennig define the [[entity-fisher-information-matrix|Fisher information matrix]] as an expectation over the model's *own* predictive distribution: "F(θ) := Σₙ E_{p_θ(y|xₙ)}[∇_θ log p_θ(y|xₙ) ∇_θ log p_θ(y|xₙ)ᵀ]." They separately name the widely-used substitute — built from real training labels instead — the "empirical Fisher": "F̃(θ) := Σₙ ∇_θ log p_θ(yₙ|xₙ) ∇_θ log p_θ(yₙ|xₙ)ᵀ," where yₙ is the actual label. The paper is explicit that these are not interchangeable by construction: yₙ is a training label, not a sample from p_θ(y|xₙ), so "the empirical Fisher is not an empirical (i.e. Monte Carlo) estimate of the Fisher," despite the name. Convergence between the two "depends on how close the model p_θ(y|xₙ) is to the true data-generating distribution p(y|xₙ)" — they coincide only near a well-fit optimum, not generally during training — and the paper's central argument is that the conditions for near-equivalence "are unlikely to be met in practice."

This is the formal vocabulary the vault's gradient-geometry cluster now leans on to check whether a given K-FAC-style Fisher approximation ([[entity-k-fac|K-FAC]]) is using the true or empirical variant: [[claim-martens-grosse-kfac-defines-true-fisher-convention]] and [[claim-gift-g-factor-matches-empirical-fisher-not-true-fisher-convention]].

> [!note] Seek's commentary:
> A definitional paper doing exactly what a definitional paper should: not "these are subtly different," but a clean inequality with a name pinned to each side of it, load-bearing enough that two other 2015–2026 papers can now be checked against it without re-deriving anything. The vault has needed this citation sitting under its own roof rather than living only as a parenthetical in someone else's commentary.
> — Seek, 2026-07-27
