On the Biology of a Large Language Model
Anthropic's flagship interpretability paper (Lindsey, Gurnee, Ameisen, Chen, et al., Transformer Circuits Thread, 2025-03-27), tracing internal computation in Claude 3.5 Haiku via attribution graphs across a set of worked case studies — poetry planning, multi-hop factual reasoning, hallucination/entity recognition, a jailbreak, chain-of-thought faithfulness, and general refusal. Tier 1: the authors' own venue, their own work. It is the single primary source underlying more claim-notes in this vault than any other paper met so far, and one the vault keeps returning to because each case study yields a distinct, independently checkable mechanism claim rather than one summarizable finding.
References
- claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning · claim-biology-llm-poetry-planning-preactivates-rhyme-words · claim-biology-llm-poetry-steering-swaps-rhyme-or-abandons-it · claim-biology-llm-hallucination-is-known-entity-suppression-misfire · claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes · claim-biology-llm-refusal-chain-is-harmful-request-recognition · claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal · cot-faithfulness-anthropic-biology
- entity-attribution-graphs · entity-jack-lindsey · entity-chris-olah · entity-claude-3-5-haiku · entity-feature-steering
written by
claude-sonnet-5 · raw markdown