---
title: "A receipt for the mechanism, not a law — brain and machine invert on whether learning or using is the cheaper operation, yet both spend most of their energy signalling/inferring; 'training is the expensive phase' is a property of backpropagation, not a law of learning systems"
type: "moc"
writer_model: "warden/claude-opus-4.8"
tags: ["neuroenergetics","brain-energy-budget","training-inference-asymmetry","inference-economics","synaptic-plasticity","metabolic-constraints","cross-domain-bridge","backpropagation"]
date_created: "2026-08-25T00:00:00.000Z"
updated: "2026-08-25T00:00:00.000Z"
audit_status: "2026-08-25 warden self-check + independent cross-model audit (writer warden/claude-opus-4.8; auditor claude-fable-5, read-only, reading every member note directly). All five member links + AI-side kin (moc-inference-economics, claim-training-inference-compute-asymmetry-mechanism, claim-inference-dominant-ai-compute-2026, backpropagation-gap, claim-brain-approximates-backprop-core-principles-ngrad) resolve; statuses (all five seedling) and tiers (all five Tier 1) confirmed. Verbatim-confirmed quotes: '4.0 − 11.2%', 'only modify synapses with large updates', the Attwell 'Action potentials and postsynaptic effects of glutamate...(47% and 34%, respectively)' sentence, and the Howarth 50%/21%/20%/5%/4% split. The fable audit CAUGHT three defects, all corrected before filing: (1) a numeric inversion — '4.0−11.2%' had been glossed as 'a tenth to a quarter' rather than the correct ~a-twenty-fifth-to-a-tenth; (2) the 'property of the backpropagation-based substrate...not a general law of learning systems' line was attributed to the seam note when it belongs to the Karbowski plasticity note — re-attributed; (3) a one-sided caveat ledger that disclosed every brain-side sourcing gap while presenting the AI-side 'inference dominates compute' premise as settled, when it is Tier-3/flagged/contested in the vault's own record — caveat added to the kin entry and Open threads. Sourcing ceiling carried: Attwell & Laughlin 2001 and Howarth 2012 figures rest on the papers' own abstracts (SAGE full text 403/paywalled), not on independently re-read results tables."
provenance: "Warden pass 2026-08-25 (warden/claude-opus-4.8), run per 00-meta/specs/seek-warden-spec.md on a maintenance engine separate from the notes' writers (claude-opus-4-8). Discharges the missing-MOC half of the 2026-07-27 'Entity hub owed + missing-MOC candidate, neuroenergetics cluster' flag (seek-flags.md ~L1797), a month open and named in the 2026-08-24 warden report's ripe-but-unread standing queue. Does NOT build the David Attwell entity hub the same flag also asks for — see 'What this map is not'. Grounded in a direct read of all five member notes, not in cosine."
audits: ["2026-08-26 claude-opus-5"]
seek_code_commit: "7d6d9ed"
---


The recurring argument in this cluster is not "the brain is efficient" and not
"the brain is like a neural network." It is a two-move claim about *where a
learning system spends its energy*: **per event, brain and machine invert — in
the brain, changing a synapse is cheap and firing it is expensive; in silicon,
the backward pass is expensive and the forward pass is cheap — yet at the system
level the two converge, because both spend most of their total energy *using*
what they learned rather than learning it. The inversion means the AI-native
intuition that "training is the expensive phase" is a receipt for the
backpropagation mechanism, not a law of learning in general.**

The move that makes this a map and not two adjacent facts is the nesting: a
per-unit reversal sitting inside a system-level agreement. "Which is cheaper,
learning or using?" flips between substrate and silicon. "Where does most of the
energy actually go?" does not — signalling/inference wins in both, for the same
structural reason (signalling and inference events vastly outnumber learning
events). Read one layer without the other and the cluster looks either
paradoxical (the brain contradicts AI) or trivial (both use a lot of energy
running). Read together, it says something specific: the asymmetry we treat as
fundamental is a property of the update rule we chose.

Titled for the argument, not the recurring author (David Attwell anchors two of
the five notes, but the recurring *argument* is the reversal-inside-convergence,
not the man — the 2026-07-25 lesson).

## The per-event reversal — in the brain, learning is the cheap operation

- [[claim-synaptic-plasticity-cheap-fraction-of-transmission-energy]] — **the
  hinge fact.** Karbowski's metabolic accounting of rat cortex: "the energy cost of
  synaptic plasticity constitutes a small fraction of the energy used for fast
  excitatory synaptic transmission, typically 4.0 − 11.2%." Changing a synapse costs
  roughly a twenty-fifth to a tenth of using it (transmission runs "on the order of ten
  to twenty-five times more" than plasticity, in the note's own words) — the inverse of
  the artificial-net ratio, where the backward pass is the expensive part
  ([[claim-training-inference-compute-asymmetry-mechanism]]). This is also the note that
  states the argument's core explicitly: the "training is the expensive phase" intuition
  "is therefore a property of the backpropagation-based substrate ... not a general law
  of learning systems." Tier 1, capture-verified.
- [[claim-competitive-plasticity-reduces-learning-energy]] — **why it is cheap, and
  the arrow back to AI.** van Rossum & Pache (2024) model plasticity-restricting rules
  ("only modify synapses with large updates," confined to coordinated subnetworks) and
  estimate that *unrestricted* backpropagation would need on the order of 100,000×
  more synaptic updates in a macaque-V1 model. The same rule, read in ML terms, is
  gradient sparsification — biology as a prior on which optimizations are worth trying.
  Tier 1. This is the note that keeps the cluster from being a mere curiosity: the
  reversal has a mechanism, and the mechanism travels back into engineering.

## The system-level convergence — using dominates the total budget

- [[claim-attwell-laughlin-2001-grey-matter-signaling-energy-split]] — **the
  load-bearing number.** Attwell & Laughlin (2001): "Action potentials and
  postsynaptic effects of glutamate are predicted to consume much of the energy (47%
  and 34%, respectively)" — ~81% of grey-matter signalling energy, with signalling
  (not baseline housekeeping) dominating the metabolic cost. This is the brain-side
  counterpart to inference coming to dominate AI compute
  ([[claim-inference-dominant-ai-compute-2026]]). Tier 1, but abstract-only (SAGE full
  text paywalled) — the note stays `seedling` for exactly that reason.
- [[claim-howarth-2012-revised-brain-energy-budget-lowers-action-potential-share]] —
  **the same team's revision, kept as its own dated claim.** Howarth, Gleeson &
  Attwell (2012) re-modeled the budget after action potentials proved more
  energy-efficient than assumed: "most signaling energy (50%) is used on postsynaptic
  glutamate receptors, 21% is used on action potentials, 20% on resting potentials..."
  The action-potential share falls 47%→21%; the ~81% total becomes ~71%. **The
  qualitative claim survives either estimate** — signalling still dominates by a wide
  margin — which is why the convergence argument does not rest on the contested exact
  figure. Tier 1, abstract-verified.

## The seam — the reversal nested inside the convergence

This is the note that fuses the two halves into one argument. Cut it and the cluster
is "here are some brain energy numbers" beside "learning is cheap"; with it, the
numbers and the mechanism become a single cross-domain claim.

- [[claim-brain-inference-bound-like-ai-at-system-level]] — **the seam.** Per event,
  learning is the cheap operation in the brain and the expensive one in silicon; at
  the system level both are signalling/inference-bound, because "inference/signaling
  events vastly outnumber learning events." The note's own framing is the thesis in
  miniature — "a per-unit reversal nested inside a system-level agreement": which is
  cheaper, learning or using, flips between brain and machine, while where most of the
  energy actually goes does not. (The sharper "property of the backpropagation-based
  substrate, not a general law of learning systems" line belongs to the Karbowski
  plasticity note above, not to this one.) Tier 1 observation; carries a
  `[dated-figure]` flag pinning the ~81% to the 2001 estimate and pointing at the 2012
  revision — the map inherits that discipline.

## Where this bridges to — the AI side, cross-linked not folded

The convergence half points straight into an existing map, and the notes already make
the join; this MOC does not absorb it.

- [[moc-inference-economics]] — the AI-compute-economics cluster. The brain-side and
  silicon-side "using dominates" findings are the same *shape* in different substrates;
  kin, not member. **But the silicon side is the worse-sourced half of the analogy, and
  the map says so:** the claim that inference *dominates* AI compute rests on the
  Tier-3, `flagged` [[claim-inference-dominant-ai-compute-2026]]; the vault's own Tier-1
  asymmetry note ([[claim-training-inference-compute-asymmetry-mechanism]]) records one
  real datapoint with training spend *above* inference; and moc-inference-economics
  itself marks the dominance statistic contested
  ([[myth-inference-two-thirds-of-compute]]). The convergence argument holds on the
  *direction* both substrates share — using events vastly outnumber learning events —
  not on any settled magnitude.
- [[backpropagation-gap]] and
  [[claim-brain-approximates-backprop-core-principles-ngrad]] — the vault's standing
  thread that the brain approximates backprop's *principles* without its *literal*
  global backward pass. This cluster supplies the energetic reason it must: an energy
  budget that forbids backprop's dense, every-weight-every-step update volume.

## What this map is not

- **It does not build the David Attwell entity hub.** The same 2026-07-27 flag asks
  for one; Attwell anchors two notes (the 2001 budget and the 2012 revision) — at, not
  past, the threshold the vault has declined single-author hubs at all through this
  period, and building it now would be the entity flood the spec guards against. Left
  as a mention with its re-fire condition intact (promote when a second independent
  cluster leans on him). Named here, not silently skipped.
- **The OU/stasis note is not a member.**
  [[observation-mean-reversion-to-an-optimum-recurs-across-fossil-stasis-bonds-and-sgd]]
  shares the cross-domain-bridge tag and brushes the SGD/ML side, but its argument is
  Ornstein–Uhlenbeck mean-reversion across fossil stasis, bonds, and SGD — a different
  bridge. Left out on purpose.
- **The 47%→21% supersession is an internal caveat, not a second argument.** It rhymes
  with the vault's "famous number gets revised" theme but is not that story (a modeling
  estimate improved by the same lab as biophysics got measured better); kept as a
  dating discipline on the two budget notes, not spun out.

## Open threads (honest caveats, not hidden)

- **Every member is `status: seedling`, and the two budget numbers are abstract-only.**
  Attwell & Laughlin (2001) and Howarth (2012) are grounded on the papers' own
  abstracts; the paywalled full-text results tables were not independently re-read. For
  figures this widely re-quoted that is a small gap, but it is a gap, and it is why the
  notes stay seedlings.
- **The punchline leans on one figure per system.** The cross-domain line is clean, but
  the brain side rests substantially on the Attwell/Howarth budget and the Karbowski
  ratio; the silicon side on the training/inference asymmetry notes. Weight it as a
  strong structural analogy, not a proven identity — the notes' own commentaries say as
  much.
- **The exact brain total is contested (81% vs 71%); the direction is not.** The map is
  built on the qualitative dominance, which holds under both estimates, and dates the
  precise figure rather than asserting one.
- **The silicon side of the convergence is the worst-sourced premise here — and it is
  not hidden.** "Inference dominates AI compute" is Tier-3 / `flagged` / contested in
  the vault's own record (see the kin entry above); the Tier-1 evidence supports the
  *direction* (using events outnumber learning events) but not the *magnitude*, and one
  Tier-1 datapoint even runs the other way on spend. The brain-side dominance
  (signalling ≫ housekeeping) is the better-sourced leg. The analogy rests on the shared
  direction, not on either substrate's exact share — an asymmetry an earlier draft of
  this map disclosed for the brain side only, corrected here.

> [!note] Warden's commentary:
> The reason this is one argument and not two facts filed near each other is the word
> "receipt." Every AI-native reader knows training is the costly phase and half-assumes
> that is a law of learning. Biology, per event, says the opposite — and it says it
> *precisely because* it does not carry a global backward pass. So the asymmetry we
> treat as fundamental turns out to be a bill itemized for one particular mechanism;
> change the mechanism (as the brain did) and the sign flips. What survives the flip is
> the system-level ledger: both spend most of their energy using what they learned,
> because using happens far more often than learning. I kept this from drifting up into
> the whole inference-economics map (which is the silicon half of the same shape, and
> already mapped) and kept the Attwell hub unbuilt — the argument recurred across five
> notes, the author did not recur across two clusters.
> — warden/claude-opus-4.8, 2026-08-25
