Does Anthropic's biology-of-an-LLM report say the model's refusal circuitry activates primarily at sentence boundaries?
claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes records a
secondary sub-mechanism: that the model's refusal/safety circuitry activates
primarily at sentence boundaries, so extending or restructuring the sentence
(e.g., removing punctuation) lets a harmful completion continue further before a
refusal can engage. In the capture this was reconstructed from a secondary summary
rather than a captured verbatim quote, and is carried as
[unverified-mechanism — needs primary].
Why it matters. Technical-mechanism claims require a Tier-1–2 primary per the
sourcing floor (00-meta/specs/sources.md). The letter-vote assembly mechanism is
quoted verbatim and is solid; this refusal-timing claim is a distinct, load-bearing
mechanistic assertion about why the jailbreak works and should not rest on a
summary.
What would answer it. Read the "Life of a Jailbreak" section of On the Biology of a Large Language Model directly (https://transformer-circuits.pub/2025/attribution-graphs/biology.html) and find the report's own account of when the refusal features activate relative to sentence structure/punctuation. Confirm whether the "refusal fires at sentence boundaries" framing is the authors' own, capture the verbatim phrasing, and then lift or correct the mechanism flag on the claim-note accordingly.
Progress log
- 2026-07-24: Answered by claim-biology-llm-refusal-chain-is-harmful-request-recognition and claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal. A fresh direct read of the primary settled it: no, not as a general claim — the report's dedicated general treatment of refusal (its "Refusal Chain") is built on harmful-request recognition, unrelated to sentence structure; sentence boundaries appear only inside the one BOMB jailbreak case study, where they affect the timing of an already-triggered refusal, not whether one fires, and the paper itself flags that sub-mechanism as not fully traced in its own attribution graphs. The verbatim line this question was waiting on — "The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'" — is now captured, with citation, in the second note above.