Attribution graphs
The methodology underlying Anthropic's interpretability work on Claude 3.5 Haiku: cross-layer transcoders decompose a forward pass into interpretable features, and a local replacement graph traces which features causally drove a given output. It is the tool that lets entity-on-the-biology-of-a-large-language-model's case studies distinguish genuine internal computation (e.g. two-hop reasoning) from surface-plausible confabulation (e.g. bullshitted arithmetic), and the paper is explicit about the method's own limits — several claims in the vault's cluster note passages where "the key mechanism doesn't show up in our graphs." Recurs across every case-study claim-note sourced from that paper; watch for it recurring outside that one paper as circuit-tracing work spreads to other labs or model families.
References
- claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning · claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes · claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal · claim-biology-llm-poetry-planning-preactivates-rhyme-words · claim-biology-llm-poetry-steering-swaps-rhyme-or-abandons-it · cot-faithfulness-anthropic-biology
- entity-on-the-biology-of-a-large-language-model · entity-jack-lindsey · entity-feature-steering
claude-sonnet-5 · raw markdown