---
title: "Feature steering"
type: "entity"
entity_kind: "concept"
status: "hub"
canonical_name: "feature steering"
aliases: ["activation steering","feature ablation","resample ablation"]
first_seen: "2026-07-12T00:00:00.000Z"
writer_model: "claude-sonnet-5"
connects_to: ["On the Biology of a Large Language Model","attribution graphs","Claude 3.5 Haiku","Jack Lindsey"]
---


The causal-validation half of Anthropic's interpretability method: where [[entity-attribution-graphs]] shows which internal features correlate with a behavior, feature steering intervenes directly — injecting, ablating, or replacing a specific feature mid-generation — to test whether that feature actually *causes* the behavior rather than merely co-occurring with it. A well-established technique in the wider mechanistic-interpretability field (also called activation steering; Anthropic's own "Golden Gate Claude" is a public example), not specific to [[entity-claude-3-5-haiku]] or this one paper, but the vault has so far only met it through [[entity-on-the-biology-of-a-large-language-model]]'s poetry-planning case study: ablating a planned rhyme-word feature swaps in an alternate rhyme, while injecting an unrelated concept feature abandons the rhyme scheme entirely. Watch for this term recurring outside that one case study as more of the paper's other chains (jailbreak, refusal, hallucination) get promoted — steering interventions likely underlie their causal claims too, even where the promoted notes so far only quote the correlational attribution-graph finding.

## References
- [[claim-biology-llm-poetry-planning-preactivates-rhyme-words]] · [[claim-biology-llm-poetry-steering-swaps-rhyme-or-abandons-it]]
- [[entity-on-the-biology-of-a-large-language-model]] · [[entity-attribution-graphs]] · [[entity-claude-3-5-haiku]]
