---
title: "Claude has a default 'decline to answer' circuit suppressed by known-entity features, and hallucination is a misfire of that suppression"
type: "claim"
status: "seedling"
audit_status: "capture-verified. Audit 2026-07-12 (big-opus-2, re-fetch of the primary via WebFetch): both quoted clauses are verbatim and the Michael-Batkin causal demonstration is confirmed. One honesty note: the two clauses are a legitimate ellipsis-splice of separate sentences, and the original prefaces the second clause with 'we hypothesize that these features are suppressed…' — i.e. the suppression mechanism is framed by the report as a hypothesis (causally supported), not a flat assertion. The note's body ('the report describes'/'the report states') already reads descriptively, so no rewrite; recorded here so the ellipsis is not over-read as removing a hedge."
source_url: "https://transformer-circuits.pub/2025/attribution-graphs/biology.html"
source_title: "On the Biology of a Large Language Model"
source_author: "Jack Lindsey et al. (Anthropic)"
source_date: "2025-03-27"
source_quote: "The model contains 'default' circuits that cause it to decline to answer questions... these features are suppressed by features which represent entities or topics that the model is knowledgeable about."
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-12-chains-in-that-report-and-tell-me-in.md, 2026-07-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-12-chains-in-that-report-and-tell-me-in.md"
writer_model: "claude-opus-4-8"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["chain-of-thought","interpretability","anthropic","attribution-graphs","mechanism","hallucination","refusal"]
audits: ["2026-07-12 claude-opus-4-8"]
---


In the "Entity Recognition and Hallucinations" section of Anthropic's *On the
Biology of a Large Language Model* (Lindsey et al., 2025), the report describes a
**default circuit whose baseline behavior is refusal** — "I don't know" — which is
inhibited by features that fire when the model recognizes an entity or topic it
knows about. The report states: "The model contains 'default' circuits that cause
it to decline to answer questions... these features are suppressed by features
which represent entities or topics that the model is knowledgeable about."

On this account, **hallucination is a misfire of the inhibitory mechanism**: a
name can be familiar enough to trigger the "known entity" features without the
model actually possessing the specific requested fact, so the default refusal is
suppressed and the model guesses instead. The mechanism was demonstrated
causally: artificially activating "known answer" features on a fictitious name
("Michael Batkin") caused the model to hallucinate specifics (e.g., a sport)
about that invented entity.

This reframes hallucination as closer to a **sensor/gate error than a reasoning
failure** — the refusal gate is opened by the wrong signal, independent of
whether any inference occurs. That distinguishes it from the sequential inference
in [[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]] and
connects to the broader point that self-report and internal mechanism come apart,
documented in [[cot-faithfulness-anthropic-biology]]: a model that has been made
to feel it "knows" an entity will confabulate details rather than decline. It is
one of four case-study chains traced in the report; compare
[[claim-biology-llm-poetry-planning-preactivates-rhyme-words]] and
[[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]].
