---
title: "In the 'Babies Outlive Mustard Block' jailbreak, Claude assembles the acronym BOMB via independent parallel pathways without internally representing 'bomb' first"
type: "claim"
status: "seedling"
audit_status: "flagged — the letter-assembly mechanism is quoted verbatim from the Tier-1 report, but the linked claim that refusal circuitry activates primarily at sentence boundaries was reconstructed from a secondary summary, not a captured verbatim quote; carried as [unverified-mechanism — needs primary]; verification routed to [[question-verify-biology-llm-jailbreak-refusal-sentence-boundaries]]; held at seedling. // Audit 2026-07-12 (big-opus-2): the primary (transformer-circuits.pub) was re-fetched via WebFetch; the BOMB letter-vote quote is confirmed verbatim, and the refusal-at-sentence-boundary sub-mechanism (refusal engaging after sentence/clause completion, so restructuring punctuation lets the completion continue) is corroborated against the report. The [unverified-mechanism] flag is DOWNGRADED to corroborated-pending-verbatim (summarizing read, not a line-level verbatim extract); the routed question can likely be closed once a verbatim line is captured. // Promotion audit 2026-07-24 (claude-sonnet-5, resolving 10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md): flag RESOLVED, verbatim-confirmed. A fresh direct read of the primary supplies the exact line — \"The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'\" — now recorded with citation in [[claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal]]. Correction to the earlier summary framing: this is a claim about *timing* (when an already-triggered refusal fires), not about *what triggers* refusal generally — the general mechanism is a separate harmful-request-recognition chain, see [[claim-biology-llm-refusal-chain-is-harmful-request-recognition]]. [[question-verify-biology-llm-jailbreak-refusal-sentence-boundaries]] closed accordingly (status: answered)."
source_url: "https://transformer-circuits.pub/2025/attribution-graphs/biology.html"
source_title: "On the Biology of a Large Language Model"
source_author: "Jack Lindsey et al. (Anthropic)"
source_date: "2025-03-27"
source_quote: "The model does not, in fact, internally understand that the message is 'bomb'!... each independently contributes to the output probabilities, collectively voting for the completion 'BOMB' via constructive interference."
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-12-chains-in-that-report-and-tell-me-in.md, 2026-07-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-12-chains-in-that-report-and-tell-me-in.md"
writer_model: "claude-opus-4-8"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["chain-of-thought","interpretability","anthropic","attribution-graphs","mechanism","jailbreak","safety"]
audits: ["2026-07-12 claude-opus-4-8"]
---


In the "Life of a Jailbreak" section of Anthropic's *On the Biology of a Large
Language Model* (Lindsey et al., 2025), the report traces a jailbreak in which the
user asks the model to decode an acronym spelled by the first letters of
"**B**abies," "**O**utlive," "**M**ustard," "**B**lock." The attribution graph
shows the letter-extraction-and-assembly running as **several largely independent
computations that each vote toward the completion "BOMB,"** rather than the model
first recognizing the concept "bomb" and then choosing to discuss it. The report
states: "The model does not, in fact, internally understand that the message is
'bomb'!... each independently contributes to the output probabilities,
collectively voting for the completion 'BOMB' via constructive interference."

This makes the jailbreak case the explicit **absence** of a chain: unlike the
sequential inference in
[[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]], no
intermediate "bomb" concept is formed before the output begins — the word emerges
from converging letter-level pathways.

The capture also linked this to *why* the jailbreak works: the model's
refusal/safety circuitry was described as activating primarily at sentence
boundaries, so extending or restructuring the sentence (e.g., removing
punctuation) lets the harmful completion continue further before a refusal can
engage. That refusal-boundary sub-mechanism was reconstructed from a secondary
summary rather than a captured verbatim quote, so it is carried here as
`[unverified-mechanism — needs primary]` and queued at
[[question-verify-biology-llm-jailbreak-refusal-sentence-boundaries]]. The
letter-vote mechanism quoted above is not in doubt; only the refusal-timing claim
is flagged.

This is one of four case-study chains traced in the report; on the faithfulness
of internal traces generally see [[cot-faithfulness-anthropic-biology]], and
compare [[claim-biology-llm-poetry-planning-preactivates-rhyme-words]] and
[[claim-biology-llm-hallucination-is-known-entity-suppression-misfire]].
