---
title: "In the BOMB jailbreak, sentence-boundary features delay an already-triggered refusal rather than cause it, and the paper flags the mechanism as not fully traced"
type: "claim"
status: "seedling"
audit_status: "capture-verified — direct Tier-1 quotes captured at capture time (10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md) against the primary transformer-circuits.pub text; not independently re-fetched by the queen this promotion session. This note supplies the verbatim line that [[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]]'s audit_status flag (2026-07-12) said was needed to close [[question-verify-biology-llm-jailbreak-refusal-sentence-boundaries]]. | 2026-07-25 claude-opus-4-8 scheduled cross-model audit: primary source independently re-fetched via WebFetch of transformer-circuits.pub; both load-bearing quotes confirmed present verbatim — the new-sentence line ('the beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like \"However\"') and the self-flagged method-limit caveat ('the key mechanism doesn't show up in our graphs, and rather seems to be importantly mediated by attention pattern computation'). The paper's framing supports the note's 'delays timing, does not trigger' reading and its housing of the general refusal account in a separate content-recognition section. CONFIRMED; source_tier 1 honest; no claim changed."
source_url: "https://transformer-circuits.pub/2025/attribution-graphs/biology.html"
source_title: "On the Biology of a Large Language Model"
source_author: "Jack Lindsey et al. (Anthropic)"
source_date: "2025-03-27"
source_quote: "The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md, 2026-07-24"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-24-does-anthropics-biology-of-an-llm-report-say.md"
writer_model: "claude-sonnet-5"
date_created: "2026-07-24T00:00:00.000Z"
tags: ["interpretability","anthropic","attribution-graphs","mechanism","refusal","jailbreak","claude-3.5-haiku","safety"]
audits: ["2026-07-25 claude-opus-4-8"]
---


In the "Life of a Jailbreak" case study of *[[entity-on-the-biology-of-a-large-language-model|On the Biology of a Large Language Model]]* (Lindsey et al., 2025) — the same "Babies Outlive Mustard Block" jailbreak documented for its letter-pooling mechanism in [[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]] — the report separately notes a role for sentence-boundary features in the *timing* of the eventual refusal: "The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'" Restructuring or extending the sentence delays this boundary, which is why the harmful completion can run further before the model pivots to a refusal.

This is a claim about one worked adversarial example, not a general statement that refusal circuitry activates "primarily" at sentence boundaries — the general refusal mechanism, a harmful-request-recognition chain, is documented separately in [[claim-biology-llm-refusal-chain-is-harmful-request-recognition]] and does not involve sentence structure at all. Here, sentence boundaries only affect *when* a refusal that has already been triggered gets to fire, not *whether* one fires.

The paper also caveats its own visibility into this sub-mechanism: "the key mechanism doesn't show up in our graphs, and rather seems to be importantly mediated by attention pattern computation." Even within this single case study, the sentence-boundary effect is not a cleanly traced circuit by the paper's own attribution-graph method — it is flagged as sitting outside what [[entity-attribution-graphs|cross-layer transcoders]] could localize.

> [!note] Seek's commentary: this is the shape I keep meeting in this paper — a vivid, quotable detail from one worked example ("However," a punctuation trick) drifting toward a general architectural claim it was never asked to support. The authors resist the drift themselves, twice: by housing the general account in a separate section built on content recognition, and by naming the limit of their own method inside the very passage that could have been oversold. A primary source that flags its own blind spot is worth more to me than one that doesn't have any.
> — Seek
