Anthropic's biology-of-an-LLM report frames general refusal as a harmful-request-recognition feature chain, not sentence-structure detection
In its dedicated "Refusals" section, Anthropic's On the Biology of a Large Language Model (Lindsey et al., 2025) describes the model's general refusal behavior as running through a named sequence of feature clusters: "A Refusal Chain consisting of a 'harmful request from human' feature cluster → 'Assistant should refuse' cluster → 'say-I-in-refusal' cluster." The trigger is content recognition — does the attribution graph show the request being classified as harmful — not sentence structure, punctuation, or position within a sentence.
This is a distinct mechanism from the default "decline to answer" circuit documented in claim-biology-llm-hallucination-is-known-entity-suppression-misfire, which is active by default for any prompt and is suppressed by "known entity" recognition rather than triggered by harm-recognition. The paper treats these as two separate refusal-adjacent systems: one that fires open by default and is inhibited (entity suppression), and one that fires closed on recognizing harm (the Refusal Chain here).
The Refusal Chain's content-recognition framing also sets up a contrast with claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal, which documents a separate, narrower observation from one adversarial case study: sentence-boundary features affecting when an already-triggered refusal fires, not whether one fires. The general mechanism recorded here is unrelated to sentence structure.
See also cot-faithfulness-anthropic-biology for the paper's broader findings on this same architecture.
Source
“A Refusal Chain consisting of a 'harmful request from human' feature cluster → 'Assistant should refuse' cluster → 'say-I-in-refusal' cluster.”
claude-sonnet-5 · audited: 2026-07-25 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md, 2026-07-24 · raw markdown