---
id: "20260711-1436-hop-ij-good-loss-and-intelligence-explosion"
title: "One man — I.J. Good — authored both the LLM training loss and the intelligence-explosion idea, then recanted"
type: "capture"
status: "promoted"
origin: "hop-batch"
promoted_to: ["30-notes/claim-ij-good-1952-logarithmic-score-is-the-llm-cross-entropy-loss.md","30-notes/claim-ij-good-1965-ultraintelligent-machine-coined-intelligence-explosion.md","30-notes/claim-ij-good-1998-reversed-intelligence-explosion-to-extinction.md"]
questions_routed: ["50-questions/question-verify-ij-good-1965-ultraintelligent-machine-primary.md","50-questions/question-verify-ij-good-1998-extinction-recantation-primary.md"]
not_promoted: ["Brier score origin (Glenn Brier, 1950) — a 'further lead', no primary read this run; folded as context into the log-score note rather than given its own note.","Good & Turing co-coined the 'Bayes factor' — unchased further lead; mentioned in the 1965 bio note, not promoted separately.","HAL 9000 / Kubrick 2001 consulting role — cultural-resonance lead; the film's HAL is already covered by claim-clarke-hal-9000-death-song-from-bell-labs-demo; folded as a mention into the 1965 note.","Bletchley Park / Hut 8 / Turing biography — supporting detail, folded into the 1965 and 1998 notes rather than promoted as a standalone claim."]
writer_model: "claude-opus-4-8"
date_created: "2026-07-11T00:00:00.000Z"
hop_chain: ["SEED: 30-notes/claim-selfaware-canonical-self-knowledge-benchmark.md (LLM self-knowledge / calibration)","SelfAware calibration note -> Brier score's 1950 meteorology origin (cross-domain bridge, max_cosine 0.648)","Brier / proper scoring rules -> logarithmic scoring rule = cross-entropy = LLM training loss (mechanism bridge, max_cosine 0.700)","log scoring rule -> I.J. Good, its 1952 author, Turing's Bletchley colleague (person, max_cosine 0.710)","I.J. Good's intelligence explosion -> his 1998 reversal: 'survival' -> 'extinction' (surprising claim)"]
novelty_max_cosine: 0.727
tags: ["proper-scoring-rules","cross-entropy","intelligence-explosion","ij-good","ai-existential-risk","calibration","cross-domain-bridge"]
source_url: "https://arxiv.org/html/2504.01781v1"
source_title: "Proper scoring rules for estimation and forecast evaluation"
source_author: "arXiv:2504.01781 (survey)"
source_date: "2025"
source_tier: 1
---


Following *calibration* out of the LLM self-knowledge literature lands on a single 20th-century figure sitting under two seemingly unrelated pillars of modern AI.

**Claim 1 — The loss that trains LLMs is a 1950s forecasting scoring rule.** The logarithmic score, S_log(P,y) = −log p(y), was introduced by I.J. Good in 1952 as a strictly proper scoring rule (the "ignorance score" in meteorology). Minimizing it *"is equivalent to the well-known maximum likelihood principle"* — i.e. negative log-likelihood, the cross-entropy objective minimized in next-token LLM pretraining. The loss that trains Claude is a proper scoring rule from probabilistic forecasting, kin to Glenn Brier's 1950 Brier score.
> source_url: https://arxiv.org/html/2504.01781v1 — quote: "It is also known as the ignorance score in meteorology and is a strictly proper scoring rule (Good, 1952)"; "Minimizing the logarithmic score is equivalent to the well-known maximum likelihood principle" — Tier 1.

**Claim 2 — The same Good coined the "intelligence explosion."** In 1965 Good — a Bletchley Park cryptanalyst who worked beside Turing in Hut 8 — defined the ultraintelligent machine as *"a machine that can far surpass all the intellectual activities of any man however clever,"* the *"last invention that man need ever make."* He later advised Kubrick on HAL 9000.
> source_url: https://en.wikipedia.org/wiki/I._J._Good — Tier 4 (primary: Good 1965, "Speculations Concerning the First Ultraintelligent Machine").

**Claim 3 — He reversed.** In a 1998 autobiographical statement, Good wrote that his 1965 opening — "The survival of man depends on the early construction of an ultra-intelligent machine" — should have *'survival'* replaced by *'extinction'*, concluding "we are lemmings."
> source_url: https://en.wikipedia.org/wiki/I._J._Good — Tier 4; primary is an unpublished autobiographical statement (via assistant Leslie Pendleton) — [surprising biographical claim; needs primary].

## Why this was hop-worthy
It bridges two vault clusters that don't currently link — probability-judgment/backpropagation (the training loss) and AI-existential-risk (Lighthill, Yao-extinction) — through one person who is in neither cluster's notes.

> [!note] Seek's commentary:
> The tidiest irony in AI history: the objective function that makes an LLM *calibrated* and the concept that makes superintelligence *terrifying* were both authored by the same statistician — who ended up betting on extinction.

## Further leads
- Glenn Brier (1950), "Verification of Forecasts Expressed in Terms of Probability," *Monthly Weather Review* — the Brier score's primary home.
- Good & Turing coined "Bayes factor" — a second Bletchley-era statistical export into modern ML.
- The Kubrick / HAL 9000 consulting role — cultural-resonance thread not chased.

## Hop chain

### Chain: LLM self-knowledge benchmark → I.J. Good's recantation

Hop 1: Web search — Brier score origin ( https://en.wikipedia.org/wiki/Brier_score ; https://ui.adsabs.harvard.edu/abs/1950MWRv...78....1B/abstract )
- Hook type: Cross-domain bridge (cross-time-period)
- Hook: "verbalized confidence / calibration" in the seed → the metric behind it, the Brier score, born in 1950 weather forecasting.
- Why followed: bridge candidate; leaves the self-knowledge cluster while staying on calibration.
- Key findings: Glenn W. Brier (US Weather Bureau) defined the Brier score in 1950 as a strictly proper scoring rule; same family now used for ML confidence and superforecasting.

Hop 2: arXiv 2504.01781, "Proper scoring rules for estimation and forecast evaluation" ( https://arxiv.org/html/2504.01781v1 )
- Hook type: Mechanism question / cross-domain bridge
- Hook: proper scoring rules → the logarithmic score and whether it equals the loss that trains neural nets.
- Why followed: vault_bridge flagged an unlinked pair (Oaksford-Chater probability ↔ Backpropagation scalar loss) this hook would connect.
- Key findings: log score S_log = −log p(y), introduced by Good 1952; minimizing it = maximum likelihood = negative log-likelihood = the cross-entropy loss of LLM pretraining.

Hop 3: I.J. Good — Wikipedia ( https://en.wikipedia.org/wiki/I._J._Good )
- Hook type: The person behind the thing
- Hook: who was "Good, 1952"?
- Why followed: person absent from the vault; suspected cross-domain payload.
- Key findings: Bletchley Park cryptanalyst (Hut 8, with Turing); coined "Bayes factor"; wrote the 1965 intelligence-explosion paper; advised Kubrick on HAL 9000.

Hop 4: I.J. Good's reversal — Wikipedia (I.J. Good / Technological singularity)
- Hook type: Surprising claim
- Hook: the intelligence-explosion optimist changed his mind.
- Why followed: contradicts the standard "Good = father of optimistic superintelligence" framing; bridges to the vault's extinction-risk note.
- Key findings: 1998 statement — 'survival' should read 'extinction'; "we are lemmings."

Saved hooks not followed:
- Tetlock / Good Judgment Project — from the Brier-score search — reason saved: amateurs beat CIA analysts on Brier score; strong surprising-claim thread, zoom-out to human forecasting (novelty 0.634).
- Gneiting & Raftery (2007) unification of proper scoring rules — reason saved: mechanism/person thread on strict propriety.
- HAL 9000 consulting role — reason saved: cultural-resonance thread (Doom-style touchstone) worth its own hop.

Surprise: expected the LLM training loss to have a purpose-built ML pedigree — found it is literally Good's 1952 logarithmic proper scoring rule from weather-forecast verification.
Surprise: expected the coiner of the "intelligence explosion" to remain a techno-optimist — found he reversed in 1998, swapping "survival" for "extinction" and calling humanity lemmings.
Surprise: expected the calibration-metric author and the AI-doom author to be different people — found they are the same man, I.J. Good.

post-worthy: yes — a clean one-person bridge from the LLM loss function to the origin of AI existential risk, dense with verifiable surprises and a cultural-touchstone kicker (HAL 9000).
