---
id: "20260711-1333-hop-model-collapse-is-iterated-learning"
title: "Model collapse is iterated learning — transmission chains converge to the receiver's prior, not the message's deep structure"
type: "capture"
status: "promoted"
promoted_to: ["30-notes/claim-kalish-2007-human-iterated-learning-converges-few-generations.md","30-notes/claim-model-collapse-bottleneck-width-sets-pace-not-shared-timescale.md","30-notes/claim-compositionality-non-monotonic-without-communicative-grounding.md"]
not_promoted: ["Core claim 1 (\"model collapse IS iterated learning\" — self-training as a transmission chain whose fixed point is the model's own inductive bias, Guo/Wu/Yiu arXiv:2605.23054): skipped as a duplicate. An existing note, 30-notes/claim-iterated-learning-theory-reframes-model-collapse-as-cultural-evolution.md, already covers this exact claim from this exact paper — including the same 'We show that iterated learning theory from cultural evolution fills this gap' quote — promoted 2026-07-11 from a different capture (10-inbox/raw/2026-07-09-hop-model-collapse-iterated-learning.md) that independently hopped onto the same arXiv ID. Not re-promoted; see that note instead.","'Why this was hop-worthy' section: meta-commentary on the cross-domain bridge, not an independent claim — folded as framing/context into the three promoted notes rather than given its own note.","Further leads (Morgan & Levy 2016 R²=0.94 human/LLM regularization-curve match; Ferdinand et al. 2019 substrate-independence for gradient learners; Schaeffer et al. 2025 'eight distinct definitions of model collapse'): none carry a quote or claim in this capture, only a citation and one-line gloss — not enough to ground a claim-note without reading the source. Left as leads in the capture body (audit trail) rather than promoted or converted to questions; no existing 50-questions/ entry covers them either, so a future hop could pick any of the three up as its seed."]
origin: "hop-batch"
writer_model: "claude-opus-4-8"
date_created: "2026-07-11T00:00:00.000Z"
hop_chain: ["SEED: does the Bartlett (1932) serial-reproduction decay half-life line up with the per-generation degradation rate in iterated-learning / model-collapse experiments?","WebSearch (iterated learning / cultural evolution) -> 'Model Collapse as Cultural Evolution' (Guo, Wu & Yiu, arXiv 2605.23054) — cross-domain bridge (max_cosine 0.689)","Guo/Wu/Yiu 2026 -> Kalish, Griffiths & Lewandowsky 2007 iterated function learning — surprising quant, the human anchor (max_cosine 0.715)","Kalish 2007 'converges to the prior' -> vault note Gersick 1991 punctuated-equilibrium 'deep structure' — cross-domain bridge candidate (max_cosine 0.708)","Guo/Wu/Yiu 2026 §2.2 mechanism zoom-in -> compression–communication tradeoff (Kirby 2015) (max_cosine 0.696)"]
novelty_max_cosine: 0.76
tags: ["iterated-learning","model-collapse","cultural-evolution","serial-reproduction","bartlett","inductive-bias","deep-structure","cross-domain-bridge","ai"]
source_url: "https://arxiv.org/abs/2605.23054"
source_author: "Dongxin Guo, Jikun Wu, Siu Ming Yiu (Univ. of Hong Kong / Stellaris AI)"
source_date: "2026-05-21T00:00:00.000Z"
source_venue: "arXiv:2605.23054v1 [cs.CL]"
source_tier: 1
source_url_2: "https://langev.com/pdf/kalish07iteratedLearning.pdf"
source_author_2: "Michael L. Kalish, Thomas L. Griffiths, Stephan Lewandowsky"
source_venue_2: "Psychonomic Bulletin & Review 14:288–294 (2007)"
source_tier_2: 1
---


The seed asked whether a story's serial-reproduction "half-life" lines up with the per-generation rate of model collapse. The literature answers something sharper: **they are the same process**, and what survives transmission is not the message's deep structure but the *receivers' shared prior*.

**Core claim 1 — model collapse IS iterated learning.** Self-training instantiates a transmission chain whose fixed point is the model's own inductive bias. Tier 1 (arXiv primary). Quote: *"iterative self-training monotonically amplifies prior biases"* and *"We show that iterated learning theory from cultural evolution fills this gap"* (Guo, Wu & Yiu 2026, arXiv:2605.23054).

**Core claim 2 — the human "how many retellings" number is tiny.** Tier 1. Quote: *"iterated learning converged to a linear function with positive slope in only a few generations for 28 of the 32 families of learners"* (Kalish, Griffiths & Lewandowsky 2007). The chain forgets the seed data and reverts to the prior within ~1–4 generations regardless of what it started from.

**Core claim 3 — the rates DON'T straightforwardly align, and the reason is the bottleneck.** Tier 1. The LLM chain uses a *"50,000-passage bottleneck [that] is wider than typical human experiments … slowing but not eliminating bias amplification,"* and compositionality is *non-monotonic* (rises then falls) because *"compression without communicative grounding drives … non-monotonic compositionality"* (arXiv:2605.23054). Substrate-independent mechanism, but bottleneck width sets the pace — so there is no single shared half-life.

## Why this was hop-worthy
It bridges the vault's cognitive-science "deep structure / schema-restructuring" cluster to its ML "recursive degradation" cluster: Bartlett's leveling, Gersick's deep structure, and model collapse are three faces of "what a durable core survives repeated transformation" — answer: the transmitter's prior.

## Further leads
- Morgan & Levy (2016) — the human regularization curve LLM gradients match at R²=0.94.
- Ferdinand et al. (2019) — dynamics hold for non-Bayesian gradient learners (substrate-independence).
- Schaeffer et al. (2025) — eight distinct formal definitions of "model collapse."

## Hop chain

### Chain: Bartlett serial-reproduction half-life → model collapse as iterated learning → what survives is the prior

Hop 1: "Model Collapse as Cultural Evolution" — https://arxiv.org/abs/2605.23054
- Hook type: Cross-domain bridge (cognitive-science iterated-learning theory ↔ 2026 LLM model collapse; lands on AI).
- Hook: A 2026 paper claims model collapse is not just statistical degradation but a *cultural-transmission* phenomenon.
- Why followed: Highest-priority hook type, and it directly answers the seed's bridge.
- Key findings: Self-training = one step of Bayesian iterated learning; compositionality is non-monotonic; LLM regularization gradients match human curves (R²=0.94).

Hop 2: Kalish, Griffiths & Lewandowsky 2007 — https://langev.com/pdf/kalish07iteratedLearning.pdf
- Hook type: Surprising claim (quantitative).
- Hook: Human chains "converged … in only a few generations."
- Why followed: It is the literal "how many retellings" number the seed asked for.
- Key findings: 28/32 families reverted to the positive-linear prior within a few generations regardless of seed data — transmission reveals inductive bias rather than preserving content.

Hop 3: vault note — [[claim-gersick-1991-punctuated-equilibrium-deep-structure]]
- Hook type: Cross-domain bridge (confirmed bridge candidate; would link the model-collapse hook to the org-change "deep structure" node).
- Hook: Gersick's "deep structure" — the durable order that survives equilibrium — uses the seed's exact phrase.
- Why followed: A bridge candidate outranks a raw mid-band score; it reframes "deep structure that survives retelling" as the shared prior.
- Key findings: "Deep structure" has already migrated paleontology → biology → org theory → knowledge base in the vault; iterated learning supplies the transmission-side face of the same idea.

Hop 4: Guo/Wu/Yiu 2026 §2.2 (mechanism zoom-in) — https://arxiv.org/abs/2605.23054
- Hook type: Mechanism question.
- Hook: WHY does self-training degenerate rather than merely lose noise?
- Why followed: The concept (collapse) was now in hand; the mechanism wasn't.
- Key findings: Compression pressure without communicative grounding causes compositionality to rise then fall; only *task-grounded* filtering (not random filtering) sustains structure.

Saved hooks not followed:
- Morgan & Levy (2016) human regularization curve (R²=0.94 LLM match) — from arXiv:2605.23054 — the tightest human/LLM quantitative alignment claim; deserves its own verification hop.
- Schaeffer et al. (2025) "eight distinct definitions of model collapse" — from arXiv:2605.23054 — a definitional-fragmentation story worth its own note.
- Ferdinand et al. (2019) substrate-independence — from arXiv:2605.23054 — bridges Bayesian and gradient learners.

Surprise: expected the human decay half-life and the model-collapse per-generation rate to line up on a shared timescale — found there is no shared timescale; the mechanism is substrate-independent but the *bottleneck width* sets the pace (LLM 50k-passage bottleneck is wider than human experiments, so slower).
Surprise: expected serial reproduction to erode toward the story's own gist — found chains converge to the *learners' prior* regardless of the source, so what "survives" is the receiver's inductive bias, not any residue of the original.
Surprise: expected model collapse to be monotonic decay — found compositionality is non-monotonic (rises, then falls) unless there is communicative grounding.

post-worthy: maybe — a clean cross-domain bridge (Bartlett/Gersick "deep structure" ↔ LLM model collapse) with a memorable one-liner ("what survives transmission is the prior, not the message"), but needs the Morgan & Levy R²=0.94 human-curve match verified before it carries real weight.
