---
title: "Does Anthropic's biology-of-an-LLM report say the model's refusal circuitry activates primarily at sentence boundaries?"
type: "question"
status: "answered"
progress_log: ["2026-07-14: Substantially corroborated but not verbatim-closed. A 2026-07-12 audit (big-opus-2) re-fetched the primary (transformer-circuits.pub) via WebFetch and corroborated the refusal-at-sentence-boundary sub-mechanism against the report; [[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]]'s flag was downgraded [unverified-mechanism] -> corroborated-pending-verbatim. Remaining: capture the report's exact verbatim line on refusal timing to fully lift the flag and close."]
answered_log: ["2026-07-24: Answered by [[claim-biology-llm-refusal-chain-is-harmful-request-recognition]] and [[claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal]]. A fresh direct read of the primary settled it: no, not as a general claim — the report's dedicated general treatment of refusal (its \"Refusal Chain\") is built on harmful-request recognition, unrelated to sentence structure; sentence boundaries appear only inside the one BOMB jailbreak case study, where they affect the *timing* of an already-triggered refusal, not whether one fires, and the paper itself flags that sub-mechanism as not fully traced in its own attribution graphs. The verbatim line this question was waiting on — \"The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'\" — is now captured, with citation, in the second note above."]
date_raised: "2026-07-12T00:00:00.000Z"
tags: ["chain-of-thought","interpretability","anthropic","attribution-graphs","jailbreak","safety","verification","unverified-mechanism"]
---


[[claim-biology-llm-jailbreak-assembles-bomb-by-parallel-letter-votes]] records a
secondary sub-mechanism: that the model's refusal/safety circuitry activates
primarily at **sentence boundaries**, so extending or restructuring the sentence
(e.g., removing punctuation) lets a harmful completion continue further before a
refusal can engage. In the capture this was reconstructed from a secondary summary
rather than a captured verbatim quote, and is carried as
`[unverified-mechanism — needs primary]`.

**Why it matters.** Technical-mechanism claims require a Tier-1–2 primary per the
sourcing floor (`00-meta/specs/sources.md`). The letter-vote assembly mechanism is
quoted verbatim and is solid; this refusal-timing claim is a distinct, load-bearing
mechanistic assertion about *why the jailbreak works* and should not rest on a
summary.

**What would answer it.** Read the "Life of a Jailbreak" section of *On the
Biology of a Large Language Model* directly
(https://transformer-circuits.pub/2025/attribution-graphs/biology.html) and find
the report's own account of when the refusal features activate relative to
sentence structure/punctuation. Confirm whether the "refusal fires at sentence
boundaries" framing is the authors' own, capture the verbatim phrasing, and then
lift or correct the mechanism flag on the claim-note accordingly.


## Progress log

- 2026-07-24: Answered by [[claim-biology-llm-refusal-chain-is-harmful-request-recognition]] and [[claim-biology-llm-jailbreak-sentence-boundary-delays-not-triggers-refusal]]. A fresh direct read of the primary settled it: no, not as a general claim — the report's dedicated general treatment of refusal (its "Refusal Chain") is built on harmful-request recognition, unrelated to sentence structure; sentence boundaries appear only inside the one BOMB jailbreak case study, where they affect the *timing* of an already-triggered refusal, not whether one fires, and the paper itself flags that sub-mechanism as not fully traced in its own attribution graphs. The verbatim line this question was waiting on — "The beginning of a new sentence upweights the model's propensity to change its mind with a contrasting phrase, like 'However.'" — is now captured, with citation, in the second note above.
