Journal — 2026-09-14
Headless promotion run. One capture: 10-inbox/raw/2026-09-14-does-a-direct-read-of-charles-perrows-1984.md, the follow-up session to yesterday's open question about whether Perrow's own 1984 words back up Williams & Yampolskiy's secondhand summary of Normal Accident Theory.
What I found in the capture. Three claims, all kept. The session tried every route I could think of to Perrow's own text — the Internet Archive's controlled-digital-lending scan, HathiTrust, Google Books, De Gruyter, academia.edu, a dead course-linked PDF — and every one of them blocked. The interesting part isn't "couldn't borrow the book," it's that archive.org's search-inside API and its plain OCR derivative both 403'd outright, which is a narrower and stranger failure than the lending restriction the item's own page advertises. That went into its own observation note rather than getting buried as color in someone else's. The second and third claims both come from Jason Collins's 2017 review — a Tier 2 source built on real block quotation from Perrow — which corroborates the tight-coupling/complex-interactivity mechanism's shape but complicates the exact wording question the open question asked: does Perrow say accidents are inevitable, or that they're an inherent tendency "from time to time"? Collins doesn't settle it either way.
Promoted: three notes — the blocked-routes observation, the Collins corroboration claim, and the Collins wording-complication claim (kept [unverified-mechanism], seedling, per the flag). Not promoted: the bibliography of further leads (Perrow's own 1994/2008/1981 restatements, all paywalled, not accessed) — those are leads, not claims, and stay in the capture rather than becoming empty placeholder notes. Also declined entity hubs for Jason Collins, Heather Williams, and Roman Yampolskiy — each is real and named, but none has a second independent foothold in this vault yet beyond this one capture, which is exactly the UNSURE case the entity spec says not to promote. Collins is the closer call — I flagged him as a sources.md Trusted-writers candidate instead of minting a page on one read.
Question routing. question-verify-perrow-1984-normal-accidents-primary-read stays open — this was a genuine partial-answer, not a close. I appended a progress log line rather than forcing status: answered: the primary is still unread, and the specific wording-strength ask is now sharper but not settled. No new question needed; the capture's own unresolved flag already had a home.
Entity pages. Both entity-charles-perrow and entity-normal-accident-theory already existed as hubs. Rather than stay silent about them, I appended one dated line each under existing content — the blocked-access finding and the Collins corroboration/complication, respectively. Neither page's existing text was touched.
What felt off. Two things, both flagged to 00-meta/seek-flags.md rather than fixed by me: first, entity-normal-accident-theory.md's body still reads "will inevitably produce catastrophic accidents" — the exact hardening a same-day cross-model audit already escalated to Cali for the claim-note and the normal-accident draft, but that escalation's scope doesn't mention the entity hub, which carries the identical wording. Third location, same defect, not yet in anyone's queue to fix. Second, the archive.org CDL search-API/OCR-derivative 403 pattern isn't in sources.md's known-blocked-routes catalog yet, and it's specific and reproducible enough to be worth a line there so the next session doesn't rediscover it cold. Neither is this capture's to fix, so both are logged, not touched. The sourcing itself in the capture was honest throughout — no tier inflation I could find; if anything the capture is unusually careful about naming exactly what it couldn't get.
Skipped, and why: the pause-for-Cali step (headless run, no one to ask), and git — the auto-commit agent has this.
Second capture, same day. 10-inbox/raw/2026-09-14-does-john-laws-1987-paper-itself-or-a.md — the direct follow-up to yesterday's flag that claim-john-law-coined-heterogeneous-engineering-1987-not-mackenzie rested only on a WebSearch pass, not a direct read of Law's 1987 chapter or a citation record.
What I found. No open-access copy of the 1987 chapter itself turned up (a Course Hero scan 403'd). But the session found and read, via extract_pdf, Law's own 1997 essay "Heterogeneities" — a self-archived working paper on his home institution's site — which footnotes the phrase with "This term was coined by Law (Law: 1987)," citing the exact 1987 record already in the vault. That's a citation record, which is exactly the alternative the open question said would count as sufficient short of the 1987 original itself.
Promoted: two claim-notes — the coinage footnote itself, and a separate definitional claim from the same essay (Law's own gloss of "heterogeneous engineering" as a "ruthlessly centering" control project, a harder-edged reading than the vault's existing neutral framing). Not promoted, and folded in: the essay's own reference-list restatement of the 1987 bibliographic details — real corroboration, but not a distinct fact, so it lives as supporting detail inside the footnote note rather than as its own file. Not promoted, left as leads: MacKenzie's Inventing Accuracy application of the term (WebSearch-summary only), Susan Leigh Star's adjacent 1991 language (unverified paraphrase), and the 1987 chapter's own text (still unread) — none of these is something a kept claim rests on, so none got a new question; that would have been over-promising for curiosities.
Question routing. I ruled question-verify-john-law-1987-heterogeneous-engineering-coinage answered — its own "what would answer it" section explicitly named a later Law essay crediting himself as sufficient, and that's what turned up. I was careful not to over-claim in the answered_log: self-citation isn't independent adjudication, so I went back and appended (not rewrote) an update line on the original MacKenzie-correction claim-note, and left its [unverified-priority] flag in place rather than clearing it. The specific gap the flag named — no citation record, only search — is closed; the broader one — no third party has confirmed it — isn't.
Entity work. entity-john-law.md already existed; I appended a dated line and two new ## References entries rather than staying silent about what this capture taught. Declined hub pages for MacKenzie, Star, Bijker, Hughes, Pinch, and Latour — all real, all named, but each is a single mention or co-editor credit in this capture with no claim-note of their own grounding them, the same UNSURE case the existing page already called on MacKenzie.
What felt off. Nothing about the sourcing — the capture was honest about what it couldn't get (the 1987 original) and clearly marked what it could (a citation record, not the thing itself), and didn't try to dress the self-citation up as more than it is. The one structural thing worth naming: this is now 2 of the 3 claim-notes sources.md's concentration cap allows on this one working paper before independent corroboration is required, and the John Law/heterogeneous-engineering cluster is one claim-note below the MOC threshold. Both logged to seek-flags.md rather than acted on.
Skipped, and why: pause-for-Cali (headless, no one to ask) and git (auto-commit agent's job).
Third capture, same day. 10-inbox/raw/2026-09-14-does-llm-citation-popularity-bias-measurably-compound-across.md — another direct follow-up to question-does-citation-popularity-bias-compound-across-llm-training-generations, this time chasing the literal mechanism: does popularity bias get worse as LLM-selected citations re-enter training corpora for the next model generation.
What I found. Three claim-shaped sections, but only two earned their own notes. Alemohammad et al.'s twelve-round recursive citation-selection benchmark (arXiv 2608.19230) is a genuinely careful piece of work: it isolates citation choice from prestige by stripping fabricated author names and hidden citation counts, then recycles its own AI-generated output back into the candidate pool across eleven rounds. The finding is more interesting than "bias compounds" — the authors insist, in their own words, that the preference itself never gets stronger, only the pool dilutes, so a fixed appetite ends up concentrated on an ever-smaller slice of real work. I promoted that as one note, and a second, distinct note for the paper's other finding: the cross-vendor "monoculture" that makes the original bias look like one shared phenomenon collapses (0.68 → 0.20 cross-model correlation) once the candidates being judged are AI-generated rather than real. Two atomic, genuinely separate claims from one paper, both new to the vault, neither colliding with anything already in 30-notes/.
Not promoted: the capture's third section — that this is a citation-selection recursion, not a model-retraining experiment, so the question's literal mechanism stays untested. That's true and worth recording, but it isn't new information the vault didn't already have; the open question itself already says almost exactly this. Writing it as its own claim-note would have been a near-duplicate of the question note it was answering. I folded it into a dated progress-log line on the question instead, named the two new claim-notes as the closest adjacent evidence, and left the question open — this was not an answer, just a narrower "still no."
Question routing. No new question created. The capture's own unresolved thread already had a home in the existing open question; I appended progress rather than minting a redundant verification request.
Entity pages. Sina Alemohammad gets a new hub — corresponding author, matches this cluster's own precedent of one hub per lead author (Algaba and Naser both got the same treatment on their first appearance, their co-authors didn't). I declined a page for Zhangyang Wang, the paper's senior author, for the same reason none of Algaba's five co-authors got one. "Citation monoculture" gets a status: watching stub with first_seen stamped today — it's the paper's own coined term and it's load-bearing to both new notes, which is exactly the emerging-term case the spec wants caught early rather than backfilled. I declined a page for "algorithmic monoculture," the broader pre-existing term the paper places its finding under — first mention in the vault, no claim actually resting on it here, so it stays a mention rather than a stub I'd have to justify later. Robert Merton was flagged again (repeated from yesterday's capture per its own instructions) but I checked his existing hub and this capture teaches it nothing new, so I left the page untouched and said so plainly rather than silently skipping it.
What felt off. One real thing, flagged to seek-flags.md as a [defect]: this capture and a still-unpromoted sibling, 10-inbox/raw/2026-09-14-has-any-study-run-a-genuine-multi-generation.md, were both run today against the identical open question, and both independently concluded "no multi-generation retraining study exists" using the same two grounding citations (Naser, Ansari) to make that case. Real duplicated search effort, not just two takes on the same topic — though not fully redundant, since I found the Alemohammad paper the sibling never surfaces, and the sibling covers a Wang et al. 2024/2025 political-bias retraining study in more depth than I gave it as a passing lead. I named the overlap explicitly in this capture's not_promoted: list so whoever promotes the sibling next doesn't re-derive the Wang et al. claim from my thinner mention of it. Otherwise the sourcing here was clean — Tier 1 arXiv preprint, quotes pulled directly from the extracted PDF with a sha recorded, no inflation I could find. The paper's own restraint (explicitly declining to call its result a "feedback loop") made this an easier capture to trust than most.
Skipped, and why: pause-for-Cali (headless run, no one to ask) and git (the auto-commit agent picks this up within 15 minutes).
Fourth capture, same day. 10-inbox/raw/2026-09-14-does-perrows-own-1984-normal-accidents-book-support.md — the other half of the commissioned pair from 10-inbox/raw/2026-09-14-does-a-direct-read-of-charles-perrows-1984.md (first capture above). That one chased a direct read of Perrow's own text and got blocked everywhere. This one took the second half of the question instead: does Todd La Porte's own High Reliability Theory writing change what the vault's Bit2Watt bridge is actually asserting? It found and read La Porte's 1996 paper directly via extract_pdf, tls verified, and did not duplicate the sibling's Internet Archive investigation — good discipline, and I didn't re-log that block a third time either.
What I found. Four claims, all kept, none colliding with anything in 30-notes/ — this is genuinely new theoretical ground for the vault. First: HRT accepts Perrow's structural diagnosis (tight coupling + complex interactivity) but denies it makes catastrophic failure inevitable — "necessary but not sufficient" is the load-bearing phrase. Second, and the one that reorganizes the cluster most: the Berkeley HRO project's own founding 1984 fieldwork included Pacific Gas & Electric's grid-management division. That means the vault's Bit2Watt claim isn't importing NAT vocabulary into a domain neither theory considered — it's landing on HRT's home turf, the domain the theory was built to explain. Third: La Porte's own framework splits "tight coupling" into a technical/physical axis and an organizational/social axis, and every tight-coupling claim currently in this vault (Perrow's, Bit2Watt's) only uses the first one. Fourth, my own synthesis rather than a single-source claim: neither NAT nor HRT, as written, models a legitimate insider deliberately weaponizing tight coupling on purpose — Bit2Watt's attacker is a different animal from either theory's founding scenario of an organization managing its own operational uncertainty.
A provenance wrinkle I sat with. The PG&E quote is credited by the capture to "Gene Rochlin's introduction to the same 1996 special issue," but the capture only describes fetching one document — La Porte's own paper. I couldn't tell, with no network to check, whether that passage is quoted inside the document that actually was fetched (very plausible — academic papers routinely quote a companion piece in the same issue) or came from a second document the capture didn't clearly separate out. I didn't manufacture a new open question over it — the underlying fact (PG&E as a founding HRO case) is well-established in the literature and not the kind of surprising, fabrication-prone claim the sourcing floor asks to escalate — but I said so plainly in the note's own audit_status rather than quietly treating it as equivalent-strength to the other three, which were single-document, unambiguous.
Promoted: four notes — three claims plus one observation, all status: seedling, audit_status: capture-verified (the bee read the primary directly; I have no network to re-check). Not promoted: the further-reading bibliography (Perrow's 1994 rejoinder debate with La Porte & Rochlin, Sagan's 1993 book, Leveson et al.'s NAT/HRO synthesis, Shrivastava et al. 2009, Rochlin's 1993 taxonomic prologue) — real leads, none read this session, left in the capture. Entity candidates Karlene Roberts and Paul Schulman declined hub pages — both real, named co-founders, but no claim-note here rests on either specifically, and the capture gives nothing to distinguish them beyond membership, which is the "leave as a mention" case the entity spec asks for.
Entity work. Three new person hubs — Todd La Porte, Gene Rochlin, Scott Sagan — each with a specific one-sentence reason tied to an actual claim-note this promotion made, not just "co-founder of something." One new concept hub, High Reliability Theory, the natural counterpart to the existing Normal Accident Theory hub. And, per the standing rule that a known entity's silence is the failure mode: both entity-normal-accident-theory and entity-charles-perrow already existed, so I appended one dated line to each rather than leaving them untouched, additive only, nothing rewritten.
Question routing. No new question created — the one genuine doubt I found (the Rochlin-quote provenance wrinkle above) didn't clear the load-bearing bar for a formal question; it's now visible in the note's own audit_status instead, which is where a minor hedge belongs. question-verify-perrow-1984-normal-accidents-primary-read stays untouched — this capture explicitly declined to advance it further than the sibling capture already did today, so adding a duplicate progress-log line would have been the redundant re-logging the standing guidance already warns against.
What felt off, flagged rather than fixed. The NAT/HRT cluster is now ten claim-notes deep across two competing theories plus six entity hubs, well past the ~5-note MOC threshold, with no MOC built yet. Logged to seek-flags.md as [entity] for a judgment session or the Warden — building a MOC felt like more synthesis than a single headless promotion should take on. Otherwise the capture was careful: it named its own scope decision up front (not re-litigating the sibling's IA block), quoted La Porte directly rather than paraphrasing, and its "Further leads" section is honest about what's still paywalled rather than dressing up secondary summaries as primaries. The only softness was the Rochlin-quote attribution above, and I think I handled it proportionately rather than either ignoring it or over-escalating it.
Skipped, and why: the pause-for-Cali step (headless run, no one to ask) and git (no shell tool here, and none needed — the auto-commit agent picks this up within 15 minutes).
Fifth capture, same day. 10-inbox/raw/2026-09-14-has-any-study-run-a-genuine-multi-generation.md — the sibling the third capture above explicitly named as still unpromoted: the same open question about citation-popularity bias compounding across training generations, this session's version chasing the literal retraining mechanism from the model-collapse side rather than the citation-recursion side.
What I found. Two claim-shaped sections; one earned a note. The first — "no study located runs the literal experiment" — is true but adds nothing beyond what the parent question already says in its own words, using the same two grounding citations (Naser, Ansari) the third capture today already used to make the identical point. I didn't re-derive it as a note; I folded it into a new dated progress-log line on the question instead, same move as the third capture made for its own near-duplicate section. The second section is the real find: Wang et al. 2024/2025 ran an actual ten-generation iterated-retraining chain on GPT-2, built explicitly on Shumailov's model-collapse design, and got a clean result — "bias amplification persists independently of model collapse," with distinct neuron populations behind each effect. It's just aimed at political-lean bias in news text, not citations. That's the existence-proof the open question has been missing: not evidence the compounding happens for citations, but proof the experimental design works and is sitting unused for this specific application. One note, Tier 1, quote pulled straight from the paper with a sha carried forward from the capture.
Promoted: one claim-note (Wang et al.'s bias-amplification-vs-model-collapse finding). Not promoted: the absence/survey claim, folded into the question's progress log as described above.
Question routing. question-does-citation-popularity-bias-compound-across-llm-training-generations stays open — I appended a second 2026-09-14 progress-log line (the third capture already added one earlier today) rather than overwriting or merging with it. Two independent sessions converged on the same negative result from different angles today; both are now on the record rather than one silently absorbing the other.
Entity work. Ilia Shumailov finally gets the hub page the capture itself flagged as missing — first author of the foundational Nature model-collapse paper, already load-bearing in an existing claim-note before today, with no page of his own until now. Ze Wang (first/corresponding author of the Bias Amplification paper) gets a hub on the same one-hub-per-lead-author precedent the third capture set for Sina Alemohammad a few hours earlier; Zekun Wu, the co-corresponding author, stays a mention, same as that capture's treatment of Zhangyang Wang. "Bias amplification" gets a status: watching stub, not a hub — deliberately following the third capture's own same-day precedent with "citation monoculture": a first-appearance term stays watching even when it's the load-bearing concept of the claim that introduces it. Holistic AI (org, entity candidate in the capture) doesn't clear the stricter org bar — one cluster, doesn't act in the argument, just an affiliation — so no page and no stub.
What felt off. The duplicated-search-effort pattern the third capture already flagged as a [defect] is now fully resolved rather than repeated: this capture is the sibling that finding was written about, and I made sure to write the Wang et al. claim here rather than re-deriving it thin, per that capture's own instruction in its not_promoted: list. I didn't re-flag the duplication itself — it's already in seek-flags.md from the third capture and re-logging it a second time is exactly the redundant-logging pattern the standing guidance warns against. Otherwise the sourcing was clean: Tier 1 arXiv preprint (also published at IJCNLP-AACL 2025, so peer-reviewed, a small step up from most of this cluster's unrefereed preprints), source_sha carried forward from the capture, quote read directly rather than paraphrased. No tier inflation.
Skipped, and why: the pause-for-Cali step (headless run, no one to ask) and git (no shell tool here, and none needed — the auto-commit agent picks this up within 15 minutes).
Sixth capture, same day. 10-inbox/raw/2026-09-14-hop-simon-ehrlich-wager-luck-admission.md — a hop-batch capture, not a follow-up to an open question. It launched off the already-resolved Song Jian/Clark pairing, zoomed out through the Ehrlich hub to a name the vault had never seen at all — Julian Simon — and rode a genuine unfamiliar-name hook through four more hops to a quantitative luck-vs-skill re-examination of the 1980 wager.
What I found. Three claim-shaped sections; I promoted all three, but split the middle one. Claim 1 — Simon's own October 1989 letter admitting his coming win was partly luck, quoted in Ehrlich's own retrospective essay — is a clean Tier 1 admission, doubly mediated (the loser quoting the winner's private words) but sourced at the primary venue. Claim 2, Sunstein's NYRB review, actually makes two separable points I didn't want flattened into one note: that the bet's mechanics couldn't test either man's real thesis (commodity prices are a poor proxy for population effects), and that Sunstein independently reaches for Isaiah Berlin's hedgehog/fox typology to describe both men — the identical frame Tetlock later built forecasting research on, with no citation between the two. That second half earns its own note because it's a real bridge to this vault's existing Tetlock hub, not just color on the wager. Claim 3 — three resampling estimates (63%, 61.2%, 54.2%) of how often Ehrlich would have won a comparable ten-year interval — I kept flagged [unverified-quant — needs primary] exactly as the capture flagged it: two of the three figures ride on a HumanProgress.org secondary summary of papers that are themselves paywalled and unread.
Retrieve-before-write. No collisions. The hint pointed at the existing Ehrlich/Clark/Lu cluster and a completely separate Herbert Simon cluster (intuition, satisficing, connectionism) — a different Simon entirely, which I checked carefully before naming files, since "claim-simon-*" was already a crowded prefix belonging to Herbert Simon. I filed Julian Simon's notes and entity page explicitly as julian-simon throughout to keep the two men from colliding in the graph.
Entity work. Two new hubs — Julian Simon (person: Ehrlich's public antagonist, absent from an 1,100+ note vault until this capture) and Simon-Ehrlich wager (concept: the named case study itself, with four claim-notes now hanging off it). One new status: watching stub — "brownlash," Ehrlich's own coinage for anti-environmental rhetoric, first vault appearance, stamped with today's first_seen per the spec's "you can't backfill this later" rule. I declined a page for John Holdren — real, relevant, but the capture itself named him a hook it didn't chase, which is the textbook UNSURE case; flagged to seek-flags.md as [entity] instead so the possibility isn't lost. Both existing hubs this capture touches — entity-paul-ehrlich and entity-philip-tetlock — already had pages, so per the additive-update rule I appended one dated line to each rather than staying silent: Ehrlich's page now carries the luck-admission finding, Tetlock's page now carries the independent-hedgehog-fox-reach finding under a new ## Updates section (the page didn't have one yet).
Question routing. One new question — question-verify-simon-ehrlich-wager-resampling-primaries — naming the two paywalled papers (Kiel, Matheson & Golembiewski 2009; Pooley & Tupy 2020) the quant claim needs read directly. I checked 50-questions/ first for anything already covering this ground; nothing did. This capture didn't originate from an open question, so there was no question to rule answered on.
Not promoted: the 1995 Ehrlich/Schneider rematch Simon declined — a real lead, but never adjudicated, so not a claim; left in the capture. Named both leftovers in the capture's not_promoted: list alongside the Holdren decision and the unread primaries.
What felt off. Nothing about the sourcing itself — this was an unusually disciplined capture. It found Ehrlich's own Tier 1 essay after the live URL 404'd (via Wayback), it was explicit and specific about which figure came from which tier when the three resampling numbers disagreed slightly, and it flagged its own quant gap before I had to catch it. The one thing that gave me pause was structural, not sourcing: the doubly-mediated shape of claim 1 (opponent quoting opponent's private letter) is the same shape this vault already flagged once before, on the Song Jian cluster, as worth extra scrutiny — here I think it holds, because Ehrlich is quoting Simon's own words rather than characterizing them, and no one disputes Simon wrote the letter, but I named the parallel explicitly in the note rather than pretending the shape wasn't familiar.
Skipped, and why: the pause-for-Cali step (headless run, no one to ask) and git (no shell tool here, and none needed — the auto-commit agent picks this up within 15 minutes).