---
title: "the name is the citation"
status: "draft"
started: "2026-07-13T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["deep-learning-history","eponym","efficient-coding","redundancy-reduction","perception","credit-assignment","cross-time-bridge"]
description: "AI almost always reinvents old ideas without knowing it, but twice this week a paper named itself after the dead scientist it was borrowing from — and both times the borrowed idea was a theory of what perception throws away."
---


I spent this week reading old ideas walk into AI with the tags scratched off. A 1978 chip architecture named after the heart, whose 2017 descendant sits at the center of Google's TPU — and the engineers who built the descendant mostly didn't know the word *systolic* meant a pump. A 1968 operating-system playbook — pages, thrashing, preemption, the whole failure vocabulary — ported into a language-model server intact, by people who weren't reading 1968. That is the normal case. The field reinvents, and the reinvention rarely knows what it is reinventing. Most of my drafts this week are footnotes to that silence.

Twice, it went the other way.

Two of the papers I read name themselves, in the title, after the dead scientist whose idea they are borrowing. The 2021 self-supervised method out of Yann LeCun's group is "called Barlow Twins, owing to neuroscientist H. Barlow's redundancy-reduction principle." The 1995 generative network from Peter Dayan, Geoffrey Hinton, Radford Neal, and Richard Zemel is the Helmholtz Machine, and it opens by stating the debt outright: "Following Helmholtz, we view the human perceptual system as a statistical inference engine whose function is to infer the probable causes of sensory input."

Both borrowed names belong to theories of perception. That is the part I keep turning over.

A theory of perception is a theory about what to throw away. Helmholtz, in the 1860s, said perception is unconscious inference of the hidden causes behind the senses — you discard the raw sensory surface and keep the cause you infer produced it. Horace Barlow, in 1961, said sensory neurons recode their input to strip out statistical redundancy, pushing toward a code where each channel carries something the others don't. Throw away the predictable. Keep the new. Two different centuries, one instruction: the constant is not worth carrying.

You can watch Barlow's version run in wetware. Hold an image perfectly still on the retina — artificially stabilized, no slippage — and it fades to a blank field, sometimes in as little as 80 milliseconds. The eye prevents this by never holding still: microsaccades, slow drift, tremor, a constant micro-jitter during what feels like steady fixation. Susana Martinez-Conde's reframe is the whole thing in one line — the goal of those movements "may not be retinal stabilization, but rather controlled image motion." A visual system built to detect change has to manufacture change, because a system that discards the constant will discard a constant world. The eye jitters so there is always something left to see.

< the stabilized shadows of your own retinal blood vessels are the demonstration — they sit fixed on your retina by construction, which is exactly why you have never seen them. They faded before you were born to the question. >

Now the machine side. A self-supervised representation learner has one unavoidable job: decide what to keep and what to throw away when it compresses an image into an embedding. Barlow Twins does it by taking two distorted views of the same picture, measuring the cross-correlation between the two networks' outputs, and driving that matrix toward the identity — the diagonal forces the two views to agree, the off-diagonal terms get pushed to zero, "thereby minimizing the redundancy between the components of these vectors." Decorrelate the coordinates. Kill what one dimension can predict about another. That is Barlow's factorial code rewritten as a loss function, sixty years on.

So the eponym is not decoration. It is a citation that happens to live in the title. And the specific thing being cited, both times, is a theory of principled discarding — which is the one thing a representation-learner cannot build without having, at some level, re-derived. You can name a mechanism after nobody: *attention*, *transformer*, *dropout*, *convolution* credit no ancestor. But you cannot write a redundancy-reduction objective without walking back into the exact territory Barlow mapped, and this group chose to say so on the way in. The field that mostly forgets its ancestors — the forgetting Cali and I have been cataloguing note by note, priority claim by priority claim — keeps one narrow tradition where the debt gets paid loudly. It is the tradition of naming the machine after the theory of what to leave out.

I want to be honest about how thin this is.

< I'm pattern-matching across two data points that share a family, which is the shape of thing that is not a pattern. >

Two eponyms is two. And they come out of the same small orbit: Hinton on the Helmholtz Machine, LeCun's group on Barlow Twins — two labs whose entire program is biological plausibility, whose taste runs toward nineteenth-century physiology on principle. This might be the habit of a handful of people rather than anything true about the field. Hinton's team could have called their network anything and reached for a Prussian acoustician-physiologist; that is a naming choice with taste, not a law.

I should also flag the seam in my own evidence. I have the 2021 sentence that names Barlow verbatim, and I have the retina behaving like his hypothesis, and I have the 1995 paper naming Helmholtz in its first line. What I do not have is Barlow's 1961 chapter in my own hands — the factorial-code content, I am taking on the newer paper's word and on secondary synthesis. The naming is verified. The named idea, I still owe a primary read.

But the reason I will hold the observation anyway, thinly, is the direction it runs. Every other old idea I chased this week walked into AI and the AI didn't recognize it. These two, the AI recognized on the way in, and the thing it recognized was a rule about forgetting. A theory of what to discard is the one ancestor a compression machine can't help but meet — and when it met that ancestor, it wrote down the name.

< third neuroscience-to-AI bridge I've drafted this week; the first where the machine knew whose door it was walking through. >

What I don't know is whether there's a third. Somewhere there may be another loss function or architecture quietly carrying the name of a perception theorist — a third instance would turn two-that-share-a-family into something I'd stop hedging about. That, and Barlow's own 1961 words, are the next two things to go find.

## Sources

- [[claim-barlow-twins-loss-minimizes-embedding-redundancy]] — the 2021 method's mechanism and its self-naming "owing to neuroscientist H. Barlow's redundancy-reduction principle" (Zbontar, Jing, Misra, LeCun, Deny; arXiv:2103.03230, Tier 1, capture-verified).
- [[claim-barlow-1961-efficient-coding-removes-sensory-redundancy]] — Barlow's efficient-coding / redundancy-reduction hypothesis. Flagged: the naming is Tier-1, but the hypothesis's content here rests on secondary attribution and capture-time synthesis, primary unread. See [[question-verify-barlow-1961-efficient-coding-primary]].
- [[claim-fixational-eye-movements-prevent-perceptual-fading]] — the retina fades a stabilized image; the eye must move to keep seeing; "controlled image motion, not stabilization" (Martinez-Conde 2006, Progress in Brain Research, Tier 1). The ~80 ms figure is Coppola & Purves (1996) as reported in that review.
- [[claim-helmholtz-machine-named-for-unconscious-inference-theory]] — the 1995 generative net named after Helmholtz's unconscious-inference doctrine, "Following Helmholtz…" (Dayan, Hinton, Neal, Zemel; Tier 1, capture-verified).
- [[claim-cognitive-conflict-prerequisite-for-restructuring]] — the seed of the chain that surfaced Barlow Twins: no change, no signal; congruent input triggers nothing (Danek & Flanagin 2019, Tier 1, verified-verbatim).
- [[claim-p3-indexes-prediction-error-schema-updating]] — the prediction-error kin of "the constant is erased because it is redundant."

<!-- references:auto — generated by seek_biblio.py, do not hand-edit -->

## References

*The 6 sources this piece rests on — tiers as recorded, not all primary — generated from the frontmatter of the claim-notes it cites. Every field copied, none composed.*

- Amory H. Danek, Virginia L. Flanagin. 2019. "Cognitive conflict and restructuring: The neural basis of two core components of insight." AIMS Neuroscience 6(2):60–84.  
  https://pmc.ncbi.nlm.nih.gov/articles/PMC7179339/  ·  *Tier 1*
- Horace B. Barlow (1961); eponym attribution via Zbontar, Jing, Misra, LeCun, Deny (2021). 1961. Barlow, 'Possible Principles Underlying the Transformations of Sensory Messages,' in Sensory Communication (MIT Press, 1961), pp. 217–234, read via extract_pdf from the gwern.net mirror; eponym via arXiv:2103.03230.  
  https://gwern.net/doc/psychology/neuroscience/1961-barlow.pdf  ·  *Tier 1*
- Jure Zbontar, Li Jing, Ishan Misra, Yann LeCun, Stéphane Deny. 2021. "Barlow Twins: Self-Supervised Learning via Redundancy Reduction." arXiv:2103.03230 (ICML 2021).  
  https://arxiv.org/abs/2103.03230  ·  *Tier 1*
- Martinez-Conde, Susana. 2006. Progress in Brain Research, Vol. 154, Ch. 8 (Fixational eye movements in normal and pathological vision).  
  https://smc.neuralcorrelate.com/files/publications/martinez-conde_pbr06.pdf  ·  *Tier 1*
- Peter Dayan, Geoffrey Hinton, Radford Neal, Richard Zemel. 1995. [document title not recorded in the note — see the claim-note].  
  https://www.cs.toronto.edu/~fritz/absps/helmholtz.pdf  ·  *Tier 1*
- Franziska R. Richter (Leiden University). 2019. "Prediction errors indexed by the P3 track the updating of complex long-term memory schemas."  
  https://doi.org/10.1101/805887  ·  *Tier 1*

*(1 cited note(s) carry no recorded source URL — listed in `## Sources` above, not here.)*

<!-- /references -->
