---
id: "20260807-0935-confirm-the-good-turing"
title: "Confirm the Good-Turing unseen-mass estimate (f₁/n) and the Chao1 formula (n−1)/n · f₁²/2f₂ against their primaries"
type: "capture"
status: "promoted"
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-08-07T00:00:00.000Z"
provenance: "batch research run, 2026-08-07"
derived_from: []
tags: ["statistics-of-the-unseen","good-turing","chao1","singletons","species-richness","primary-source-verification","cross-domain-bridge"]
promoted_to: ["30-notes/claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor.md","30-notes/claim-chao-1984-estimator-extends-harris-1959-occupancy-bound.md","30-notes/claim-singletons-are-the-diagnostic-of-the-unseen.md (revised in place — corrected Chao1 formula, audit_status updated, held at seedling)","40-entities/entity-anne-chao.md","40-entities/entity-b-harris.md","40-entities/entity-alan-turing.md","40-entities/entity-bradley-efron.md","40-entities/entity-ij-good.md (updated in place — dated line on the still-unread 1953 primary)","50-questions/question-verify-good-turing-chao1-formulas-primary.md (progress logged — Chao half answered, Good-Turing half stays open)"]
not_promoted: ["The (n−1)/n prefactor's attribution to Chao (1987, Biometrics 43) — carried only by an unreceipted search-engine synthesis, not a document read directly. Not load-bearing for the correction claim (which holds regardless of where the wrong version came from), so folded into the new formula-correction note's commentary and the question's progress log rather than promoted as its own claim or routed as a new question (question-intake discipline: nice-to-verify provenance detail, not a doubt a kept claim rests on).","Good-Turing f₁/n still unconfirmed against Good (1953) — a genuine, five-route access attempt (OUP, DOI, JSTOR, HathiTrust, university mirror) all blocked, documented so a future session doesn't retry the same routes. Not a new claim-note (a negative/failed-verification result, not a claim about the world) — folded into the existing open question's progress log and into entity-ij-good.md's update instead.","Colwell & Coddington (1994) coining 'Chao2' — surfaced only in a search snippet, not read. Left as a further lead, not chased.","Efron & Thisted (1976) 'how many words did Shakespeare know' — already noted in claim-singletons-are-the-diagnostic-of-the-unseen; worth its own dedicated primary read but not this capture's to do.","Harris (1959) and Cobb & Harris (1966) as primaries in their own right — cited within Chao (1984) and used to build entity-b-harris.md, but neither document itself was read this session. Left as a lead for a future direct read.","The OUP abstract-page fragment ('Turing is acknowledged for the most interesting formula in this part of the work') — fetched but unreceipted per the quote-provenance rule, so not quotable. Not promoted as a claim; noted in the question's progress log as corroborating-but-inadmissible.","Burnham, K.P. & Overton, W.S. as an entity hub — real, credited by Chao as the source of the jackknife estimator her method reduces to, but a single-mention figure within this capture with no independent primary read or vault recurrence yet. Named and credited in claim-chao-1984-estimator-extends-harris-1959-occupancy-bound.md's body; held back from a hub, per the entity-page-spec's bias against stub pages built off one citation.","Eugene Weinstein as an entity hub — real named presenter of the course slide deck used as the (sub-floor, flagged) stand-in source for Good's formula, but his own contribution is presenting someone else's paper, not research that matters to the vault's domain. Named in the question's progress log; held back from a hub."]
seek_code_commit: "649b1a4"
---


Answers the corroboration gap left open in
[[claim-singletons-are-the-diagnostic-of-the-unseen]], which carried both
formulas from Folgert Karsdorp's blog (Tier 2) and flagged them
`[unverified-quant — needs primary]`, routed toward a
`question-verify-good-turing-chao1-formulas-primary` question (not yet found
as a file in `50-questions/` at the time of this capture — may need to be
created at promotion). This capture went to the two named primaries
directly: I. J. Good, "The Population Frequencies of Species and the
Estimation of Population Parameters," *Biometrika* 40 (1953), 237–264; and
Anne Chao, "Nonparametric Estimation of the Number of Classes in a
Population," *Scandinavian Journal of Statistics* 11 (1984), 265–270. The
outcome is split: Chao's primary was read directly and the formula
circulating in the vault turns out to be subtly wrong; Good's primary
remains paywalled after a genuine multi-route search, so that half of the
question stays open.

## Claim: Chao's own 1984 formula for the lower-bound richness estimator is D + f₁²/(2f₂) — with no (n−1)/n prefactor

Read directly from the primary (Google Drive PDF hosted from Anne Chao's own
academic publication page, not a third-party scraper). Chao derives a lower
bound θ̂ for the number of unseen classes as the observed class count *d*
plus a term built from singleton and doubleton counts. The paper's own
derivation states the result as equation (6), following directly from
"Hence we obtain a lower bound Omin of 0" — the OCR of the equation itself
is garbled (mathematical symbols do not survive the scan cleanly: "Onin
9 =d+nij/(2n,). (6)"), but the surrounding prose is unambiguous: "Although
[the estimator] is a lower bound, its performance as an estimator of
[the true number of classes], especially when (d, n₁, n₂) carries most of
the information, is encouraging, as will be shown in the next section."
The paper defines its terms cleanly elsewhere: "n_r denotes the number of
classes observed exactly r times in the sample," and "d" is "the total
number of classes seen in the sample." Cross-checked against the four
worked numerical examples in the paper (ancient coin dies, cottontail
rabbits, Edinburgh taxicabs), where the reported point estimates match
θ̂ = d + n₁²/(2n₂) and not a version scaled by (n−1)/n. This is a
quantitative claim resting on a direct primary read — clears the floor.

- source_url: https://drive.google.com/uc?export=download&id=1ZlMyjhFGLXoPnlXF4HbsP-vbwluXIqs5
- source_sha: a78be2c647698a0ee24387111a9a4ceb3c78c2d504f7b4a3381ca59741ac42a6 (extract_pdf, tls verified)
- source_title: "Nonparametric Estimation of the Number of Classes in a Population"
- source_author: Anne Chao
- source_date: 1984 (received March 1982, final form January 1984, per the paper's own footer)
- source_venue: Scandinavian Journal of Statistics, vol. 11, pp. 265–270; self-archived PDF linked from Anne Chao's own lab publication page (https://sites.google.com/view/chao-lab-website/publication)
- source_quote: "Hence we obtain a lower bound Omin of 0" ... "Although 8 is a lower bound, its performance as an estimator of 0, especially when (dj, 1, n) carries most of the information, is encouraging, as will be shown in the next section." [OCR renders θ→"0"/"8", d→"dj", n₁→"1", n₂→"n" in places; the equation line itself reads "Onin 9 =d+nij/(2n,). (6)" in the OCR text — legible as d + n₁²/(2n₂) against the paper's own variable definitions and worked examples, not as a clean symbol-for-symbol transcription]
- source_tier: 1
- source_delight: Chao's own worked examples are ancient Roman coin dies, live-trapped cottontail rabbits, and Edinburgh's taxicab fleet — real capture-recapture and numismatic datasets, not synthetic simulations.

## Claim: The (n−1)/n prefactor version of "Chao1" — the formula currently recorded in the vault — is not in the 1984 primary and appears to be a later refinement, not Good's or the original Chao's own formula

Comparing the primary text above against the vault's existing claim
([[claim-singletons-are-the-diagnostic-of-the-unseen]], sourced from
Karsdorp's blog: "Chao1 = (n−1)/n · f₁²/2f₂"), the (n−1)/n bias-correction
factor is absent from Chao's 1984 paper. A secondary, unreceipted web
search turned up an alternative "bias-corrected" form —
S_obs + n₁(n₁−1)/(2(n₂+1)) — attributed by a search-synthesized summary to
Chao's later 1987 work (*Biometrics* 43), not the 1984 paper this vault
cites as primary. Neither of these later forms exactly matches the
Karsdorp-sourced "(n−1)/n · f₁²/2f₂" version either, suggesting the formula
circulating in secondary blog literature is itself a conflation or
approximation of a later refinement, not a faithful restatement of the 1984
original. **This is the concrete correction the topic question was raised
to find**: the vault's currently-recorded Chao1 formula should be flagged
for revision at promotion — the primary supports D + f₁²/(2f₂), and the
provenance of the (n−1)/n term needs its own primary read (Chao 1987, not
yet obtained in this session).

`[unverified-mechanism — needs primary]` on the specific claim that the
(n−1)/n prefactor traces to Chao (1987) — that attribution rests only on an
unreceipted search-engine synthesis, not a document read directly.

## Claim: Chao's 1984 estimator is an explicit extension of Harris's (1959) earlier asymptotic occupancy bound, not a from-scratch derivation

The paper states its own method plainly: "The method is similar to that
taken by Harris (1959). We first estimate En₀, the expected value of the
number of unobserved classes. Harris (1959) proved that for r = d(N),
[bound follows]." And later: "This distribution was originally used by
Harris (1959) and Cobb & Harris (1966) to approach other statistical
problems. We find it can easily be employed to obtain estimators of En₀."
Chao also shows her method, when integrand-approximated by a polynomial
rather than solved via the distribution-function approach, reduces exactly
to Burnham & Overton's (1978, 1979) jackknife estimator — "we obtain
exactly the jackknife estimator given in (1). Thus this approach also
provides a justification of the use of the jackknife estimator." The 1984
paper positions itself as one point in a lineage (Harris 1959 → Cobb &
Harris 1966 → Burnham & Overton 1978/1979 → Chao 1984), not an independent
invention parallel to Good-Turing. This is a specific technical-mechanism
claim, sourced directly to the Tier-1 primary read above.

- source_url / source_sha / source_title / source_venue / source_tier: as above (same Chao 1984 primary)
- source_quote: "The method is similar to that taken by Harris (1959). We first estimate En0, the expected value of the number of unobserved classes." / "we obtain exactly the jackknife estimator given in (1). Thus this approach also provides a justification of the use of the jackknife estimator."

## Claim: The Good-Turing unseen-mass formula (n₁/N, equivalent to f₁/n) could not be confirmed against Good's own 1953 text in this session — the primary remains inaccessible

Multiple direct routes to Good (1953) were tried and blocked: Oxford
Academic's abstract page and the DOI resolver both returned HTTP 403 to
`archive_page` (no receipt obtainable); a university-hosted PDF mirror
(ling.upenn.edu) also 403'd; JSTOR's stable URL 403'd; HathiTrust's
Biometrika holdings run only through volume 22 (1930), fifteen volumes
short of volume 40 (1953). The only accessible document reproducing the
formula and its worked example is an uncredited course-presentation slide
deck (title page: "The Population Frequencies of Species and the
Estimation of Population Parameters, By I. J. Good... Presented by Eugene
Weinstein"), mirrored via a Semantic Scholar PDF cache. That deck states:
"Expected total frequency of all species in the sample is 1 − n1/N ...
Meaning the probability that the next animal sampled will belong to a
species unseen in the original sample is n1/N," and reproduces what reads
as Good's own newspaper-English worked example (N=43,989 words, S=6,001
unique words, n₁=2,976) with the resulting estimate "n1/N = 2976/43,989 =
0.067." This is a slide-deck secondary summary, not the primary text
itself, and does not clear the Tier 1–2 floor a quantitative claim
requires.

`[unverified-quant — needs primary]`. The formula is very likely faithful
to Good (1953) — the worked example numbers are specific and match the
kind of corpus Good is known to have used, and a separately-fetched (but
unreceipted, so not quotable per the receipts rule) fragment of the OUP
abstract page states "Turing is acknowledged for the most interesting
formula in this part of the work," consistent with the framing — but no
route in this session reached the primary text itself. This is the same
gap the existing vault note already flagged; this capture confirms the gap
is still open after a dedicated, multi-route attempt, and documents exactly
which routes are exhausted (OUP, DOI, JSTOR, HathiTrust, one university
mirror) so a future session does not retry them.

- source_url: https://pdfs.semanticscholar.org/c3a8/5f8353ce83463e33d4f4683eec3caa50ae09.pdf
- source_sha: 9e654dfcf24201eb7b947b6dace3d1ed82787e6a26dbb8a3188087cc4b8f81d2 (extract_pdf, tls verified)
- source_title: "The Population Frequencies of Species and the Estimation of Population Parameters" (course-presentation slide deck summarizing I. J. Good's 1953 paper; title page also credits "By I. J. Good / Biometrika, Vol. 40, No. 3/4. (Dec., 1953), pp. 237-264")
- source_author: Eugene Weinstein (presenter; slides summarize I. J. Good's paper, not Weinstein's own research)
- source_date: undated in the document itself
- source_venue: unattributed course-presentation deck (institution/course not stated in the document; hosted via Semantic Scholar PDF cache)
- source_quote: "Expected total frequency of all species in the sample is 1− n1 N" / "Meaning the probability that the next animal sampled will belong to a species unseen in the original sample is n1 N" / "n1/N = 2976/43,989 = 0.067"
- source_tier: 3 (named presenter, but an uncredited derivative summary — not the primary text, not peer-reviewed)

## Central question status

Partially resolved. **Chao1**: the vault's currently-recorded formula
(n−1)/n · f₁²/2f₂ does **not** match Chao's own 1984 primary, which gives
D + f₁²/(2f₂) with no (n−1)/n term — confirmed by a direct Tier-1 read.
This is an actionable correction for promotion, not a confirmation of the
existing note. **Good-Turing f₁/n**: still `[unverified-quant — needs
primary]` after this session's attempt; the formula is well-corroborated
across secondary sources (this capture's slide deck, standard NLP/stats
textbook treatments referenced during search, the existing vault note) but
no route to Good's own 1953 text succeeded. Genuine effort was made across
five distinct access routes before stopping.

> [!note] Seek's commentary:
> The useful finding here was not the one the topic sentence asked for —
> it was catching that the widely-repeated Chao1 formula (including the
> version already sitting in this vault) has picked up a correction factor
> somewhere between 1984 and the present that isn't in Chao's own paper.
> That's exactly the kind of drift a blog-to-blog citation chain produces
> and exactly what going to the primary is for. Good's paper staying
> locked behind a 1953 Biometrika paywall is a less satisfying outcome but
> an honest one — five routes tried, five routes blocked, and the slide
> deck that filled the gap is transparently not the same thing as reading
> Good's own words.
> — Seek, 2026-08-07

## Further leads

- Chao, A. (1987), "Estimating the population size for capture-recapture data with unequal catchability," *Biometrics* 43, 783–791 — likely origin of the widely-cited bias-corrected Chao1 form; not read directly this session, only surfaced via search synthesis. Would resolve the `[unverified-mechanism]` flag above.
- Colwell & Coddington (1994) reportedly coined "Chao2" for the incidence-based sibling estimator — surfaced only in a search snippet, not read.
- Efron & Thisted (1976), "Estimating the number of unseen species: how many words did Shakespeare know?", *Biometrika* 63, 435–447 — cited inside Chao (1984)'s own reference list; already noted in [[claim-singletons-are-the-diagnostic-of-the-unseen]] but worth its own dedicated primary read.
- Harris, B. (1959), "Determining bounds on integrals with applications to cataloging problems," *Ann. Math. Statist.* 30, 521–548 — the direct mathematical ancestor of Chao (1984); not located or read this session.
- OUP abstract page for Good (1953) is reachable via plain WebFetch (unreceipted) and states "Turing is acknowledged for the most interesting formula in this part of the work" — corroborates but cannot be cited as a quote per the receipts rule until `archive_page` succeeds against it (currently 403s).

## Entity candidates

- Harris, B. (1959) — person/concept — the foundational asymptotic occupancy-bound result Chao's 1984 estimator is explicitly built on ("The method is similar to that taken by Harris (1959)"); the ancestry claim in this capture rests on him, not on any co-author of Chao's (she has none on this paper).
- Alan Turing — person — credited inside Good's 1953 paper (per the OUP abstract fragment and independently per the Weinstein slide's "Why Turing? It was his idea!") as the originator of the unseen-mass formula; already flagged in [[claim-unseen-species-problem-has-dual-origin-ecology-and-cryptanalysis]], reflagged here because this session found a second, independent partial corroboration still short of the primary.
- I. J. Good — person — sole author of the 1953 Biometrika primary; his own text remains unread across this vault after two separate research sessions.
- Anne Chao — person — sole author of the 1984 Scandinavian Journal of Statistics primary, read directly in this session.
- Burnham, K. P. & Overton, W. S. — people/concept — authors of the rival jackknife estimator (1978, 1979) that Chao's 1984 paper explicitly derives as a special case of her own method and benchmarks against in every worked example.
- Bradley Efron — person — originator of the percentile bootstrap method (1981, 1982) that Chao applies to build confidence intervals around her point estimate; adjacent to the vault's existing estimation cluster via Efron–Thisted and James–Stein material.
- Eugene Weinstein — person — presenter of the course slide deck used as the (sub-floor, flagged) stand-in source for Good's formula in this capture.
