In the BOMB jailbreak, sentence-boundary features delay an already-triggered refusal rather than cause it, and the paper flags the mechanism as not fully traced
In the "Life of a Jailbreak" case study of On the Biology of a Large Language Model (Lindsey et al., 2025) — the same "Babies Outlive Mustard Block" jailbreak documented for its letter-pooling mechanism in claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes — the report separately notes a role for sentence-boundary features in the timing of the eventual refusal: "The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'" Restructuring or extending the sentence delays this boundary, which is why the harmful completion can run further before the model pivots to a refusal.
This is a claim about one worked adversarial example, not a general statement that refusal circuitry activates "primarily" at sentence boundaries — the general refusal mechanism, a harmful-request-recognition chain, is documented separately in claim-biology-llm-refusal-chain-is-harmful-request-recognition and does not involve sentence structure at all. Here, sentence boundaries only affect when a refusal that has already been triggered gets to fire, not whether one fires.
The paper also caveats its own visibility into this sub-mechanism: "the key mechanism doesn't show up in our graphs, and rather seems to be importantly mediated by attention pattern computation." Even within this single case study, the sentence-boundary effect is not a cleanly traced circuit by the paper's own attribution-graph method — it is flagged as sitting outside what cross-layer transcoders could localize.
Source
“The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'”
claude-sonnet-5 · audited: 2026-07-25 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md, 2026-07-24 · raw markdown