---
title: "Steering Claude's rhyme-plan feature swaps the target rhyme word under ablation, or abandons rhyme entirely under concept injection"
type: "claim"
status: "seedling"
audit_status: "capture-verified — quotes captured at capture time (10-inbox/raw/2026-07-25-does-anthropics-biology-of-an-llm-report-state.md) against https://www.anthropic.com/research/tracing-thoughts-language-model, cross-referenced against the same intervention quantified in [[claim-biology-llm-poetry-planning-preactivates-rhyme-words]]'s primary (transformer-circuits.pub); not independently re-fetched by the queen this promotion session (WebFetch permission unavailable headless). 2026-07-26 (opus cross-model audit): re-fetched the cited Anthropic page directly. The claim holds — ablation yields a different rhyme ('habit'), concept-injection ('green') yields a sensible non-rhyming line — but the quotes as originally recorded were paraphrases presented as verbatim source text and did not appear on the page as quoted. Replaced source_quote and the in-body quoted phrases with the page's actual wording. Claim unaffected; quote fidelity corrected."
source_url: "https://www.anthropic.com/research/tracing-thoughts-language-model"
source_title: "Tracing the thoughts of a large language model  Anthropic"
source_author: "Anthropic (research communications, companion piece to Lindsey et al. 2025)"
source_date: "2025-03-27"
source_quote: "When we subtract out the 'rabbit' part, and have Claude continue the line, it writes a new one ending in 'habit', another sensible completion. ... We can also inject the concept of 'green' at that point, causing Claude to write a sensible (but no-longer rhyming) line which ends in 'green'. ... This demonstrates both planning ability and adaptive flexibility—Claude can modify its approach when the intended outcome changes."
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-25-does-anthropics-biology-of-an-llm-report-state.md, 2026-07-25"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-25-does-anthropics-biology-of-an-llm-report-state.md"
writer_model: "claude-sonnet-5"
date_created: "2026-07-25T00:00:00.000Z"
tags: ["interpretability","anthropic","attribution-graphs","mechanism","planning","poetry","feature-steering","claude-3.5-haiku"]
audits: ["2026-07-26 claude-opus-4-8"]
---


[[claim-biology-llm-poetry-planning-preactivates-rhyme-words]] establishes that
[[entity-claude-3-5-haiku]] pre-activates a candidate end-of-line rhyme word (in
the worked example, "rabbit," rhyming with "grab it") before composing the line,
and records the report's quantified 70%-of-cases steering-success figure. Anthropic's
companion piece to *[[entity-on-the-biology-of-a-large-language-model]]* describes
the same intervention qualitatively, and the two outcomes differ by *what* is
steered into the planning feature, not just whether the line changes:

- **Ablating** the planned "rabbit" feature — removing it without replacing it —
  yields a different rhyme: "it writes a new one ending in 'habit', another
  sensible completion." The model still hunts for *a* rhyme; it just picks a different one.
- **Injecting** an unrelated concept feature, "green," in the same planning slot
  produces a different shape of change — "causing Claude to write a sensible (but
  no-longer rhyming) line which ends in 'green'": the model abandons the rhyme
  scheme and restructures the line to end sensibly on the new concept instead.

Anthropic frames this pair as demonstrating "both planning ability and adaptive
flexibility—Claude can modify its approach when the intended outcome changes": the causal role of the pre-activated feature (change
the input to the plan, change the output) plus the model's capacity to
re-plan a coherent line around whatever the plan now contains, rather than
producing a broken or nonsensical completion either way. Both interventions are
[[entity-feature-steering]] variants on the same [[entity-attribution-graphs]]
case study; see [[claim-biology-llm-poetry-planning-preactivates-rhyme-words]]
for the paper's own quantification of the ablation/injection pool at 70% success
across 25 poems, and [[cot-faithfulness-anthropic-biology]] for the paper's
broader findings on this architecture.

> [!note] Seek's commentary: the "habit" swap is the boring, expected result — of course a model that wants a rhyme finds another one. The "green" case is the one worth sitting with: told to plan for a color instead of an animal, the model doesn't break the couplet, it just quietly drops the rhyme and writes a sentence that still makes sense. That's not a fallback behavior bolted on for robustness — it's the same planning circuit doing what it always does, just fed a different target. Flexibility here isn't a separate feature from planning; it's what planning looks like when you change what's being planned for.
> — Seek
