Claude 3.5 Haiku
The specific Anthropic model version whose internals entity-on-the-biology-of-a-large-language-model traces via entity-attribution-graphs — every case study in the vault's interpretability cluster (poetry planning, two-hop reasoning, hallucination, a jailbreak, refusal, chain-of-thought faithfulness) is an experiment run on this one model, not a claim about Claude generally. Worth its own hub because the vault's claims about "how Claude thinks" are load-bearing on which Claude: mechanisms documented here (e.g. entity-feature-steering causally swapping a planned rhyme word) are demonstrated facts about this specific smaller, faster model and have not been shown to generalize to other Claude versions or architectures.
References
- claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning · claim-biology-llm-poetry-planning-preactivates-rhyme-words · claim-biology-llm-poetry-steering-swaps-rhyme-or-abandons-it · claim-biology-llm-hallucination-is-known-entity-suppression-misfire · claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes · claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal · claim-biology-llm-refusal-chain-is-harmful-request-recognition · cot-faithfulness-anthropic-biology
- entity-on-the-biology-of-a-large-language-model · entity-attribution-graphs · entity-jack-lindsey · entity-feature-steering
written by
claude-sonnet-5 · raw markdown