Does Anthropic's biology-of-an-LLM report state that steering the planned rhyme word changed the completion in ~70% of resamples?
claim-biology-llm-poetry-planning-preactivates-rhyme-words records that
injecting an alternative planned-word feature changed the eventual rhyme word in a
majority of resampled completions, and the capture attached a specific figure of
70%. That number was reproduced from a search-engine summary of the report,
not confirmed against the report's own text, and is carried as
[unverified-quant — needs primary].
Why it matters. Quantitative claims require a Tier-1–2 primary per the
sourcing floor (00-meta/specs/sources.md), and a bare percentage lifted from a
secondary summary is exactly the kind of figure that garbles in translation. The
core planning mechanism is quoted verbatim and is not in question; only the exact
steering-success percentage is.
What would answer it. Read the "Planning in Poems" section of On the Biology of a Large Language Model directly (https://transformer-circuits.pub/2025/attribution-graphs/biology.html) and locate the steering/intervention result. Confirm (a) whether a specific percentage is stated, (b) whether it is 70% or another value, and (c) exactly what it measures (fraction of resampled completions whose rhyme word shifted to the injected target). If confirmed, record the verbatim figure and lift the quant flag; if the report gives no such number, drop the figure from the note.
ANSWERED 2026-07-25. A dedicated capture read the primary directly and confirmed the figure verbatim: "we injected two planned word features ('rabbit' and 'green') in a random sample of 25 poems, and found that the model ended its line with the injected planned word in 70% of cases." Yes to the substance — 70% is correct and Tier-1-sourced. One correction to this question's own framing: the report's denominator is "25 poems" / "70% of cases," not "resamples" — that word was this vault's paraphrase, not the report's. Flag lifted on claim-biology-llm-poetry-planning-preactivates-rhyme-words.
Progress log
- claim-biology-llm-poetry-planning-preactivates-rhyme-words — the report's own text, cross-checked across three independent WebFetch reads of transformer-circuits.pub, states: "we injected two planned word features ('rabbit' and 'green') in a random sample of 25 poems, and found that the model ended its line with the injected planned word in 70% of cases." Settles the number (70%, verbatim) but corrects the question's own framing: the report's denominator is "25 poems" / "70% of cases," not "resamples" — substantively yes, terminologically not a literal match.