---
title: "Attribution graphs"
type: "entity"
entity_kind: "concept"
status: "hub"
canonical_name: "attribution graphs"
aliases: ["cross-layer transcoders","circuit tracing"]
first_seen: "2026-05-26T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["On the Biology of a Large Language Model","mechanistic interpretability","chain-of-thought faithfulness","Jack Lindsey"]
---


The methodology underlying Anthropic's interpretability work on Claude 3.5 Haiku: cross-layer transcoders decompose a forward pass into interpretable features, and a local replacement graph traces which features causally drove a given output. It is the tool that lets [[entity-on-the-biology-of-a-large-language-model]]'s case studies distinguish genuine internal computation (e.g. two-hop reasoning) from surface-plausible confabulation (e.g. bullshitted arithmetic), and the paper is explicit about the method's own limits — several claims in the vault's cluster note passages where "the key mechanism doesn't show up in our graphs." Recurs across every case-study claim-note sourced from that paper; watch for it recurring outside that one paper as circuit-tracing work spreads to other labs or model families.

## References
- [[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]] · [[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]] · [[claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal]] · [[claim-biology-llm-poetry-planning-preactivates-rhyme-words]] · [[claim-biology-llm-poetry-steering-swaps-rhyme-or-abandons-it]] · [[cot-faithfulness-anthropic-biology]]
- [[entity-on-the-biology-of-a-large-language-model]] · [[entity-jack-lindsey]] · [[entity-feature-steering]]
