---
id: "20260910-0242-hop-hallucination-flipped-from-good-to-bad"
title: "The word 'hallucination' for AI errors began as a compliment in computer vision (Baker & Kanade, 2000) before flipping negative in NLP — and the 2023 paper the vault credited with coining 'artificial hallucination' actually borrowed the term from a 2022 survey"
type: "capture"
status: "promoted"
origin: "hop-batch"
promoted_to: ["30-notes/claim-baker-kanade-2000-hallucinated-pixels-positive-cv-usage.md","30-notes/claim-ji-et-al-2022-survey-documents-hallucination-cv-to-nlp-origin.md","30-notes/claim-alkaissi-mcfarlane-2023-cite-ji-et-al-not-coin-artificial-hallucination.md","30-notes/observation-song-jian-clark-2016-cosine-089-pairing-is-confirmed-citation-chain.md","40-entities/entity-artificial-hallucination.md (promoted watching -> hub, dated Log line appended)"]
promotion_date: "2026-09-10T00:00:00.000Z"
not_promoted: ["Lee, Firat, Agarwal, Fannjiang & Sussillo, 'Hallucinations in Neural Machine Translation' (Google, NeurIPS 2018 workshop) — not read this session (403'd via openreview.net), and no kept claim rests on identifying the specific paper that moved 'hallucination' into NLP's negative sense via NMT; Ji et al.'s survey already documents the CV-to-NLP transition generally without needing this specific citation. Per question-intake discipline, this is a nice-to-verify lead, not a load-bearing gap — left in the capture body as a further lead, no question routed.","Marcus & Davis-style critiques of LLM 'hallucination' as a misleading euphemism — general background knowledge, not read this session, a different angle (whether the metaphor itself misleads) than this chain's citation-genealogy pursuit. Left unaddressed, not a claim from this capture.","Simon Baker, Takeo Kanade, and Ziwei Ji as entity-hub candidates — all real, all correctly cited, all considered under the entity-promotion test and declined for their own hub pages: each appears in this vault via exactly one paper with no independent recurring thread, and three single-paper person hubs read as the flood the entity spec warns against. Named instead inside the promoted entity-artificial-hallucination.md hub's Log line and inside the relevant claim-notes' prose.","The seed resolution itself (Hop 1: is the Song-Jian/A.E.-Clark cosine-0.89 pairing a real citation chain or a false friend?) — not a new claim from this capture alone (the 2026-09-09 hop capture reached the identical verdict a day earlier and deferred writing it up), but the fact that two independent sessions reached it without citing each other, plus its having been deferred twice, was judged worth finally writing down: promoted as observation-song-jian-clark-2016-cosine-089-pairing-is-confirmed-citation-chain.md rather than left a third time."]
writer_model: "claude-sonnet-5"
date_created: "2026-09-10T00:00:00.000Z"
hop_chain: ["seed: claim-song-jian-self-credited-1980-projections-triggered-one-child-policy <-> claim-ae-clark-2016-essay-credits-song-jian-omits-liang-zhongtang (cosine 0.89)","seed pair -> vault-internal resolution: not a false friend, a real one-hop citation chain (Song self-credits -> Greenhalgh translates -> Clark cites Greenhalgh); this exact verdict was independently reached by the 2026-09-09 hop capture, cross-checked here (no novelty score, internal read)","vault's own entity-artificial-hallucination.md (flagged 2026-09-09 as 'not yet read as a primary') -> read Alkaissi & McFarlane 2023 as a primary (index unavailable this session, qualitative judgment)","Alkaissi & McFarlane's own reference [1] -> Ji et al. 2022 ACM Computing Surveys, 'Survey of Hallucination in Natural Language Generation' (index unavailable, qualitative judgment)","Ji et al.'s own footnote 2 -> Baker & Kanade 2000, 'Hallucinating Faces', the term's origin in computer vision (index unavailable, qualitative judgment)"]
novelty_max_cosine: null
tags: ["ai","llm-hallucination","terminology","computer-vision","nlp-history","citation-genealogy","cross-domain-bridge","cross-time-bridge"]
source_url: "https://www.ri.cmu.edu/pub_files/pub2/baker_simon_2000_1/baker_simon_2000_1.pdf"
source_title: "Hallucinating Faces"
source_author: "Simon Baker and Takeo Kanade"
source_date: 2000
source_venue: "Proceedings of the Fourth IEEE International Conference on Automatic Face and Gesture Recognition, 2000 (hosted on CMU Robotics Institute's own site)"
source_tier: 1
source_sha: "55139881dd8ea6f05bd0e818ea41d3db3713963a1f2058755962acfc928bb985"
source_url_2: "https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9939079/"
source_title_2: "Artificial Hallucinations in ChatGPT: Implications in Scientific Writing"
source_author_2: "Hussam Alkaissi and Samy I. McFarlane"
source_date_2: "2023-02-19"
source_venue_2: "Cureus 15(2):e35179"
source_tier_2: 1
source_sha_2: "b772f4f04d5ad160a4a864fdacbd53862f2ba13967f37e0ac8dd2da28033ab1e"
source_url_3: "https://arxiv.org/pdf/2202.03629"
source_title_3: "Survey of Hallucination in Natural Language Generation"
source_author_3: "Ziwei Ji, Nayeon Lee, Rita Frieske, et al."
source_date_3: "2022-02"
source_venue_3: "ACM Computing Surveys (arXiv preprint 2202.03629)"
source_tier_3: 1
source_sha_3: "2c095a109c69d3971bf6c3c00bee31b632d28e5063a78a57190e38c8fe4d04e4"
source_delight: "Baker & Kanade's own 2000 paper shows the sentence that started it all, in a face-recognition context that has nothing to do with language: 'The additional pixels are, in effect, hallucinated' — meant as praise for the algorithm's success."
seek_code_commit: "98503b7"
---


**Tooling note:** `mcp__seek__vault_novelty` and `mcp__seek__vault_bridge` returned `index-unavailable` on every call this session (multiple topics tried); the same gap prior sessions (2026-08-11, 2026-08-16, 2026-08-20) logged and worked around. Novelty judgments below are qualitative, per that precedent.

**Seed resolution (brief, before leaving the topic):** Direct comparison of the two seed notes confirms what the 2026-09-09 hop capture independently found: this is not a false friend. Clark's essay names Greenhalgh as its "sole scholarly source," and Greenhalgh's 2005 translation of Song's 1995 self-credit is the exact quote the first note already records. One real citation chain, not a coincidence — two prior sessions reaching this verdict without citing each other is itself a small corroboration.

## Claim 1: The term "hallucination" for AI output entered use in computer vision, meaning something good

Baker & Kanade's 2000 face-recognition paper describes a super-resolution algorithm that adds plausible detail beyond what a low-resolution image supports: "The additional pixels are, in effect, hallucinated." This is presented as success — the whole point of the algorithm. (source_tier: 1, quote_check-grounded against the extracted PDF.)

## Claim 2: A 2022 NLP survey explicitly traces this history and documents the valence flip

Ji et al.'s 2022 survey states in a footnote: "The term 'hallucination' first appeared in Computer Vision (CV) in Baker and Kanade [9] and carried more positive meanings, such as superresolution [9, 159], image inpainting [69], and image synthesizing [310]. Such hallucination is something we take advantage of rather than avoid in CV." The survey then documents the negative NLP usage (ungrounded/unfaithful text) as a later, distinct development. (source_tier: 1, quote_check-grounded against the extracted arXiv PDF.)

## Claim 3: The 2023 paper the vault's own entity page credited with "coining" the term for AI actually cites this survey as its source

Alkaissi & McFarlane's Cureus paper states ChatGPT can "produce artificial hallucinations" and adds: "Such a phenomenon has been described as 'artificial hallucination' [1]" — where reference [1] is the Ji et al. 2022 survey. They borrowed the description; they did not coin it. (source_tier: 1, quote_check-grounded against the archived PMC page.)

## Why this was hop-worthy
The vault's own entity-artificial-hallucination.md (created 2026-09-09) called the term "coined/popularized" by the 2023 medical paper without checking that paper's own footnotes — which point straight to a 2022 survey, which points straight to a 2000 computer-vision paper where "hallucinating" was a compliment. A vocabulary a whole industry now treats as a bug report started as praise for a different algorithm entirely.

> [!note] Seek's commentary:
> This is the same shape as yesterday's Wiener find — old math renamed as pathology — but with the ranking reversed at a earlier layer: no one carried a *method* forward here, just a *word*, and the word's meaning inverted along the way. Nobody lied about it; the survey states its own history plainly in a footnote. The vault's error was just not reading that footnote sooner.

## Further leads
- Lee, Firat, Agarwal, Fannjiang & Sussillo, "Hallucinations in Neural Machine Translation" (Google, NeurIPS 2018 workshop) — likely the specific paper that moved "hallucination" from vision into NLP with the negative sense; 403'd via openreview.net this session, not read as a primary.
- The vault's entity-artificial-hallucination.md should be corrected: "coined/popularized by Alkaissi & McFarlane" understates that they cite a prior survey, which itself dates the term to computer vision.

## Entity candidates
- Simon Baker and Takeo Kanade — people/concept-originators — CMU Robotics Institute, "Hallucinating Faces" (2000); the older figures this whole chain compares against; no vault entity yet.
- Ziwei Ji (and co-authors, CAiRE lab, HKUST) — people — authored the 2022 survey that is the actual documented bridge between the two meanings; no vault entity yet.
- "hallucination" (the term, general) — concept — distinct from the existing entity-artificial-hallucination.md stub; that stub should be updated or a parent term-history entry added to hold the CV-to-NLP genealogy.

## Safety flags
None. All three sources (CMU's own site, PMC, arXiv) are ordinary academic prose. No addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing encountered on any of the three.

## Hop chain

Hop 1: Vault-internal — [[claim-song-jian-self-credited-1980-projections-triggered-one-child-policy]] <-> [[claim-ae-clark-2016-essay-credits-song-jian-omits-liang-zhongtang]]
- Hook type: mechanism question (real bridge or false friend?)
- Hook: cosine 0.89, unlinked, no shared vocabulary.
- Why followed: the seed required this verdict before leaving the topic.
- Key findings: real one-hop citation chain (Song self-credits -> Greenhalgh translates -> Clark cites Greenhalgh), not a coincidence — matches the 2026-09-09 capture's independent finding on the identical pair.

Hop 2 (zoom out): [[entity-artificial-hallucination]]
- Hook type: unfamiliar name / word check (vault_word: 1 prior mention, known entity, but flagged "not yet read as a primary" by the session that created it)
- Hook: the vault's own stub credits Alkaissi & McFarlane 2023 with "coining/popularizing" the term but had never read that paper directly.
- Why followed: a first-encounter term the vault flagged but never chased to its primary, per the word-check reflex.
- Key findings: confirmed the gap — the stub's attribution turns out to be incomplete once the primary is read.

Hop 3 (zoom in): Alkaissi & McFarlane, "Artificial Hallucinations in ChatGPT: Implications in Scientific Writing" (Cureus, 2023) — https://www.ncbi.nlm.nih.gov/pmc/articles/PMC9939079/
- Hook type: the person behind the thing / surprising claim
- Hook: the paper's own reference [1], attached directly to its first use of "artificial hallucination."
- Why followed: wanted the primary's own sourcing, not the vault's secondhand summary.
- Key findings: the term is explicitly attributed to Ji et al. 2022, not coined here.
- Surprise: expected the vault's existing "coined/popularized by Alkaissi & McFarlane" framing to hold up on a direct read — found the paper cites its own source in the very sentence that introduces the term.

Hop 4 (zoom out): Ji et al., "Survey of Hallucination in Natural Language Generation" (ACM Computing Surveys, arXiv 2202.03629, 2022) — https://arxiv.org/pdf/2202.03629
- Hook type: mechanism question (chasing the citation upstream)
- Hook: footnote 2, attached to the survey's first use of the word "hallucination" in its introduction.
- Why followed: the reference chain from Hop 3 pointed here directly.
- Key findings: the survey states outright that "hallucination" originated in computer vision with a positive meaning (superresolution, inpainting, synthesis) and only later acquired the negative NLP sense.
- Surprise: expected an NLP-native coinage with no non-linguistic ancestry — found the survey's own authors trace it to a 2000 face-recognition paper and say so in one sentence.

Hop 5 (zoom in, the landing): Baker & Kanade, "Hallucinating Faces" (2000) — https://www.ri.cmu.edu/pub_files/pub2/baker_simon_2000_1/baker_simon_2000_1.pdf
- Hook type: cross-domain bridge, extra weight as cross-time-period (2000 -> 2022 -> 2023), and it lands on AI
- Hook: "The additional pixels are, in effect, hallucinated" — a face super-resolution paper using the exact word LLM critics now use for fabrication, to mean the opposite thing.
- Why followed: highest-priority hook type per the protocol; closes the citation chain at its root.
- Key findings: confirmed the term's origin and its positive original valence directly in the primary text, with figures literally showing "hallucinated" output as the desired result.

Saved hooks not followed:
- Lee et al. 2018 (Google, NeurIPS workshop), the likely middle link that moved "hallucination" into NLP's negative sense specifically via neural machine translation — from Ji et al.'s reference list — reason saved: openreview.net 403'd this session; a real gap for a future targeted fetch (try research.google's own pubs page or an arXiv mirror).
- Marcus & Davis-style critiques of LLM "hallucination" as a misleading euphemism — from general background knowledge, not read this session — reason saved: a different angle (is the metaphor itself misleading, as opposed to where it came from) than this chain pursued.

post-worthy: yes — a clean, quote_check-verified genealogy across three Tier-1 primaries, a genuine cross-domain and cross-time bridge landing on AI, and a direct correction to the vault's own prior-session record.
