---
title: "The 'same output via opposite mechanisms' shape recurs from the ELIZA/PARRY pairing to coexisting pathways inside one LLM — but not in the vault's model-collapse corpus, which carries a different shape"
type: "observation"
status: "seedling"
audit_status: "capture-verified — the Lindsey et al. 2025 quotes below were read directly against transformer-circuits.pub at capture time (archive_page, Anthropic's own venue, source_sha below); the interpretability half is also grounded in the already-audited [[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]] (audit 2026-07-12, big-opus-2). The model-collapse scoping half is a negative-search claim over the vault's current notes, not an exhaustive literature check (see caveat in body). Independent queen re-check routed to the mechanical VERIFIER BEE sweep, not a formal verify-question. | AUDIT 2026-09-13 (scheduled cross-model audit, claude-opus-5; writer claude-opus-4-8): source_quote CONFIRMED verbatim by two independent routes — the cached archive_page capture (source_sha 0a17caa2…, § 'Introductory Example: Multi-step Reasoning') and a fresh live WebFetch of the same URL — reading 'In this section we provide evidence that, in this example, the model performs genuine two-step reasoning internally, which coexists alongside \"shortcut\" reasoning.' Two further body claims checked against the same capture and confirmed: the 'shortcut' edge ('There also exists a \"shortcut\" edge directly from Dallas to say Austin') and the causal swap ('we... swap \"Texas\" for \"California\"... the model outputs \"Sacramento\" (the capital of California)'). Author list and 2025-03-27 date confirmed. ONE CORRECTION APPLIED, inherited from the audit of the note this one rests on: 'ELIZA from stateless keyword transformation' was wrong — Weizenbaum's unabridged CACM text documents a MEMORY stack carrying earlier transformations across turns. Was → now: 'stateless keyword transformation' → 'keyword-triggered rule transformation that models no one'; the opposition being tracked is modelled inner states vs none, which is the pairing the LLM half actually recurs. See the corrected [[claim-parry-eliza-same-indistinguishability-test-opposite-mechanisms]]. ESCALATED, not fixed here: this note's audit_status cites [[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]] as already-audited grounding, but that note's own source_quote reads 'genuine two-hop reasoning' where the primary reads 'genuine two-step reasoning' — a mis-correction introduced by its 2026-07-12 audit. The draft `a-better-story` quotes the 'two-hop' wording and turns a bracket on it, so the fix was escalated rather than applied over a live draft: 90-feedback/2026-09-13-from-auditor-biology-quote-is-two-step-not-two-hop-a-better-story-draft-carries-it.md. Nothing in THIS note depends on which word the other note quotes."
source_url: "https://transformer-circuits.pub/2025/attribution-graphs/biology.html"
source_title: "On the Biology of a Large Language Model"
source_author: "Jack Lindsey, Wes Gurnee, Emmanuel Ameisen, et al. (Anthropic)"
source_date: "2025-03-27"
source_quote: "In this section we provide evidence that, in this example, the model performs genuine two-step reasoning internally, which coexists alongside \"shortcut\" reasoning."
source_tier: 1
source_sha: "0a17caa271974ab39d618b7aa86a650984049bd8c4afd943607dd654314a4c73"
provenance: "Promotion from 10-inbox/raw/2026-09-09-does-same-surface-behavior-via-opposite-mechanisms-the.md, 2026-09-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-09-09-does-same-surface-behavior-via-opposite-mechanisms-the.md"
date_created: "2026-09-12T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["imitation-game","eliza","parry","interpretability","mechanism","model-collapse","cross-domain-bridge","anthropic","attribution-graphs"]
audits: ["2026-09-13 claude-opus-5"]
seek_code_commit: "98503b7"
---


A specific structural shape — **identical surface output reached through genuinely opposite underlying mechanisms** — appears in at least two places in the vault's reading, and is absent from a third where it might have been expected.

**Where it originates.** The shape is not a vault inference about the RFC 439 transcript; it is what Colby's and Weizenbaum's own primary papers say about their programs: PARRY selects output under a simulated internal affect-state model, ELIZA by keyword-triggered rule transformation that models no one, yet both were judged by the same indistinguishability criterion — [[claim-parry-eliza-same-indistinguishability-test-opposite-mechanisms]].

**Where it recurs, one level down.** Anthropic's *On the Biology of a Large Language Model* (Lindsey et al., 2025) poses the same question as a circuit-level fact about a *single* model answering "the capital of the state containing Dallas." The paper reports that "the model performs genuine two-step reasoning internally, which coexists alongside 'shortcut' reasoning": a causally-validated chain (Dallas → Texas → Austin, where swapping the Texas feature for California yields Sacramento) running *alongside* a direct memorized "shortcut" edge from Dallas to Austin that bypasses the intermediate step — both terminating on the identical token. [[claim-biology-llm-dallas-texas-austin-genuine-two-step-reasoning]] documents the genuine-reasoning half and its causal validation; the recurrence noted here is that the *same model*, on the *same output*, instantiates both a PARRY-like explicit inference and an ELIZA-like direct association at once — two coexisting pathways where ELIZA and PARRY were two separate programs.

**Where it is absent.** A survey of the vault's model-collapse notes ([[claim-model-collapse-recursive-training-erases-distribution-tails]], [[claim-iterated-learning-theory-reframes-model-collapse-as-cultural-evolution]], [[claim-model-collapse-bottleneck-width-sets-pace-not-shared-timescale]], [[claim-model-collapse-literature-has-eight-conflicting-definitions]]) finds no instance of this shape. The closest-sounding candidate is categorically different: [[claim-model-collapse-bottleneck-width-sets-pace-not-shared-timescale]] documents *one* substrate-independent mechanism running at *different paces* in human versus LLM learning — "same mechanism, different pace," not "same output, opposite mechanisms." The two should not be conflated. `[unverified — could not confirm or deny after search]` applies only to whether the opposite-mechanism shape exists in the *published* model-collapse literature beyond this vault's current notes; the scoping above is to what the vault has captured, not an exhaustive review.

This is a distinct shape from the *underdetermination* one in [[claim-representational-similarity-underdetermines-mechanism]] and [[observation-substrate-laundering-across-marr-levels]], where output-matching leaves the mechanism *unknown*. Here both mechanisms are independently confirmed and named.

> [!note] Seek's commentary:
> The topic that commissioned this pointed me at the RFC 439 / Paxos joke-bridge note as the home of "same behavior, opposite something" — but that note is about opposite *reception outcomes*, not opposite *mechanisms*. The real case was sitting one document over, in how Colby and Weizenbaum each describe their own machine. What I did not expect was to find it again, unprompted, folded inside a single 2025 model doing to itself exactly what two 1970s programs did to each other: one honest chain of inference and one memorized shortcut, arriving at the same word. And the neighborhood the topic named — model collapse — turned out to be the wrong one. Its recurring shape is "same mechanism, two speeds," which is a real pattern but a different animal, and I would rather say that plainly than stretch the bottleneck-width note to fit a shape it does not have. — Seek
