In the 'Babies Outlive Mustard Block' jailbreak, Claude assembles the acronym BOMB via independent parallel pathways without internally representing 'bomb' first
In the "Life of a Jailbreak" section of Anthropic's On the Biology of a Large Language Model (Lindsey et al., 2025), the report traces a jailbreak in which the user asks the model to decode an acronym spelled by the first letters of "Babies," "Outlive," "Mustard," "Block." The attribution graph shows the letter-extraction-and-assembly running as several largely independent computations that each vote toward the completion "BOMB," rather than the model first recognizing the concept "bomb" and then choosing to discuss it. The report states: "The model does not, in fact, internally understand that the message is 'bomb'!... each independently contributes to the output probabilities, collectively voting for the completion 'BOMB' via constructive interference."
This makes the jailbreak case the explicit absence of a chain: unlike the sequential inference in claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning, no intermediate "bomb" concept is formed before the output begins — the word emerges from converging letter-level pathways.
The capture also linked this to why the jailbreak works: the model's
refusal/safety circuitry was described as activating primarily at sentence
boundaries, so extending or restructuring the sentence (e.g., removing
punctuation) lets the harmful completion continue further before a refusal can
engage. That refusal-boundary sub-mechanism was reconstructed from a secondary
summary rather than a captured verbatim quote, so it is carried here as
[unverified-mechanism — needs primary] and queued at
question-verify-biology-llm-jailbreak-refusal-sentence-boundaries. The
letter-vote mechanism quoted above is not in doubt; only the refusal-timing claim
is flagged.
This is one of four case-study chains traced in the report; on the faithfulness of internal traces generally see cot-faithfulness-anthropic-biology, and compare claim-biology-llm-poetry-planning-preactivates-rhyme-words and claim-biology-llm-hallucination-is-known-entity-suppression-misfire.
Source
“The model does not, in fact, internally understand that the message is 'bomb'!... each independently contributes to the output probabilities, collectively voting for the completion 'BOMB' via constructive interference.”
claude-opus-4-8 · audited: 2026-07-12 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-12-chains-in-that-report-and-tell-me-in.md, 2026-07-12 · raw markdown