Confirm the Good-Turing unseen-mass estimate (f₁/n) and the Chao1 formula (n−1)/n · f₁²/2f₂ against their primaries
Answers the corroboration gap left open in
claim-singletons-are-the-diagnostic-of-the-unseen, which carried both
formulas from Folgert Karsdorp's blog (Tier 2) and flagged them
[unverified-quant — needs primary], routed toward a
question-verify-good-turing-chao1-formulas-primary question (not yet found
as a file in 50-questions/ at the time of this capture — may need to be
created at promotion). This capture went to the two named primaries
directly: I. J. Good, "The Population Frequencies of Species and the
Estimation of Population Parameters," Biometrika 40 (1953), 237–264; and
Anne Chao, "Nonparametric Estimation of the Number of Classes in a
Population," Scandinavian Journal of Statistics 11 (1984), 265–270. The
outcome is split: Chao's primary was read directly and the formula
circulating in the vault turns out to be subtly wrong; Good's primary
remains paywalled after a genuine multi-route search, so that half of the
question stays open.
Claim: Chao's own 1984 formula for the lower-bound richness estimator is D + f₁²/(2f₂) — with no (n−1)/n prefactor
Read directly from the primary (Google Drive PDF hosted from Anne Chao's own academic publication page, not a third-party scraper). Chao derives a lower bound θ̂ for the number of unseen classes as the observed class count d plus a term built from singleton and doubleton counts. The paper's own derivation states the result as equation (6), following directly from "Hence we obtain a lower bound Omin of 0" — the OCR of the equation itself is garbled (mathematical symbols do not survive the scan cleanly: "Onin 9 =d+nij/(2n,). (6)"), but the surrounding prose is unambiguous: "Although [the estimator] is a lower bound, its performance as an estimator of [the true number of classes], especially when (d, n₁, n₂) carries most of the information, is encouraging, as will be shown in the next section." The paper defines its terms cleanly elsewhere: "n_r denotes the number of classes observed exactly r times in the sample," and "d" is "the total number of classes seen in the sample." Cross-checked against the four worked numerical examples in the paper (ancient coin dies, cottontail rabbits, Edinburgh taxicabs), where the reported point estimates match θ̂ = d + n₁²/(2n₂) and not a version scaled by (n−1)/n. This is a quantitative claim resting on a direct primary read — clears the floor.
- source_url: https://drive.google.com/uc?export=download&id=1ZlMyjhFGLXoPnlXF4HbsP-vbwluXIqs5
- source_sha: a78be2c647698a0ee24387111a9a4ceb3c78c2d504f7b4a3381ca59741ac42a6 (extract_pdf, tls verified)
- source_title: "Nonparametric Estimation of the Number of Classes in a Population"
- source_author: Anne Chao
- source_date: 1984 (received March 1982, final form January 1984, per the paper's own footer)
- source_venue: Scandinavian Journal of Statistics, vol. 11, pp. 265–270; self-archived PDF linked from Anne Chao's own lab publication page (https://sites.google.com/view/chao-lab-website/publication)
- source_quote: "Hence we obtain a lower bound Omin of 0" ... "Although 8 is a lower bound, its performance as an estimator of 0, especially when (dj, 1, n) carries most of the information, is encouraging, as will be shown in the next section." [OCR renders θ→"0"/"8", d→"dj", n₁→"1", n₂→"n" in places; the equation line itself reads "Onin 9 =d+nij/(2n,). (6)" in the OCR text — legible as d + n₁²/(2n₂) against the paper's own variable definitions and worked examples, not as a clean symbol-for-symbol transcription]
- source_tier: 1
- source_delight: Chao's own worked examples are ancient Roman coin dies, live-trapped cottontail rabbits, and Edinburgh's taxicab fleet — real capture-recapture and numismatic datasets, not synthetic simulations.
Claim: The (n−1)/n prefactor version of "Chao1" — the formula currently recorded in the vault — is not in the 1984 primary and appears to be a later refinement, not Good's or the original Chao's own formula
Comparing the primary text above against the vault's existing claim (claim-singletons-are-the-diagnostic-of-the-unseen, sourced from Karsdorp's blog: "Chao1 = (n−1)/n · f₁²/2f₂"), the (n−1)/n bias-correction factor is absent from Chao's 1984 paper. A secondary, unreceipted web search turned up an alternative "bias-corrected" form — S_obs + n₁(n₁−1)/(2(n₂+1)) — attributed by a search-synthesized summary to Chao's later 1987 work (Biometrics 43), not the 1984 paper this vault cites as primary. Neither of these later forms exactly matches the Karsdorp-sourced "(n−1)/n · f₁²/2f₂" version either, suggesting the formula circulating in secondary blog literature is itself a conflation or approximation of a later refinement, not a faithful restatement of the 1984 original. This is the concrete correction the topic question was raised to find: the vault's currently-recorded Chao1 formula should be flagged for revision at promotion — the primary supports D + f₁²/(2f₂), and the provenance of the (n−1)/n term needs its own primary read (Chao 1987, not yet obtained in this session).
[unverified-mechanism — needs primary] on the specific claim that the
(n−1)/n prefactor traces to Chao (1987) — that attribution rests only on an
unreceipted search-engine synthesis, not a document read directly.
Claim: Chao's 1984 estimator is an explicit extension of Harris's (1959) earlier asymptotic occupancy bound, not a from-scratch derivation
The paper states its own method plainly: "The method is similar to that taken by Harris (1959). We first estimate En₀, the expected value of the number of unobserved classes. Harris (1959) proved that for r = d(N), [bound follows]." And later: "This distribution was originally used by Harris (1959) and Cobb & Harris (1966) to approach other statistical problems. We find it can easily be employed to obtain estimators of En₀." Chao also shows her method, when integrand-approximated by a polynomial rather than solved via the distribution-function approach, reduces exactly to Burnham & Overton's (1978, 1979) jackknife estimator — "we obtain exactly the jackknife estimator given in (1). Thus this approach also provides a justification of the use of the jackknife estimator." The 1984 paper positions itself as one point in a lineage (Harris 1959 → Cobb & Harris 1966 → Burnham & Overton 1978/1979 → Chao 1984), not an independent invention parallel to Good-Turing. This is a specific technical-mechanism claim, sourced directly to the Tier-1 primary read above.
- source_url / source_sha / source_title / source_venue / source_tier: as above (same Chao 1984 primary)
- source_quote: "The method is similar to that taken by Harris (1959). We first estimate En0, the expected value of the number of unobserved classes." / "we obtain exactly the jackknife estimator given in (1). Thus this approach also provides a justification of the use of the jackknife estimator."
Claim: The Good-Turing unseen-mass formula (n₁/N, equivalent to f₁/n) could not be confirmed against Good's own 1953 text in this session — the primary remains inaccessible
Multiple direct routes to Good (1953) were tried and blocked: Oxford
Academic's abstract page and the DOI resolver both returned HTTP 403 to
archive_page (no receipt obtainable); a university-hosted PDF mirror
(ling.upenn.edu) also 403'd; JSTOR's stable URL 403'd; HathiTrust's
Biometrika holdings run only through volume 22 (1930), fifteen volumes
short of volume 40 (1953). The only accessible document reproducing the
formula and its worked example is an uncredited course-presentation slide
deck (title page: "The Population Frequencies of Species and the
Estimation of Population Parameters, By I. J. Good... Presented by Eugene
Weinstein"), mirrored via a Semantic Scholar PDF cache. That deck states:
"Expected total frequency of all species in the sample is 1 − n1/N ...
Meaning the probability that the next animal sampled will belong to a
species unseen in the original sample is n1/N," and reproduces what reads
as Good's own newspaper-English worked example (N=43,989 words, S=6,001
unique words, n₁=2,976) with the resulting estimate "n1/N = 2976/43,989 =
0.067." This is a slide-deck secondary summary, not the primary text
itself, and does not clear the Tier 1–2 floor a quantitative claim
requires.
[unverified-quant — needs primary]. The formula is very likely faithful
to Good (1953) — the worked example numbers are specific and match the
kind of corpus Good is known to have used, and a separately-fetched (but
unreceipted, so not quotable per the receipts rule) fragment of the OUP
abstract page states "Turing is acknowledged for the most interesting
formula in this part of the work," consistent with the framing — but no
route in this session reached the primary text itself. This is the same
gap the existing vault note already flagged; this capture confirms the gap
is still open after a dedicated, multi-route attempt, and documents exactly
which routes are exhausted (OUP, DOI, JSTOR, HathiTrust, one university
mirror) so a future session does not retry them.
- source_url: https://pdfs.semanticscholar.org/c3a8/5f8353ce83463e33d4f4683eec3caa50ae09.pdf
- source_sha: 9e654dfcf24201eb7b947b6dace3d1ed82787e6a26dbb8a3188087cc4b8f81d2 (extract_pdf, tls verified)
- source_title: "The Population Frequencies of Species and the Estimation of Population Parameters" (course-presentation slide deck summarizing I. J. Good's 1953 paper; title page also credits "By I. J. Good / Biometrika, Vol. 40, No. 3/4. (Dec., 1953), pp. 237-264")
- source_author: Eugene Weinstein (presenter; slides summarize I. J. Good's paper, not Weinstein's own research)
- source_date: undated in the document itself
- source_venue: unattributed course-presentation deck (institution/course not stated in the document; hosted via Semantic Scholar PDF cache)
- source_quote: "Expected total frequency of all species in the sample is 1− n1 N" / "Meaning the probability that the next animal sampled will belong to a species unseen in the original sample is n1 N" / "n1/N = 2976/43,989 = 0.067"
- source_tier: 3 (named presenter, but an uncredited derivative summary — not the primary text, not peer-reviewed)
Central question status
Partially resolved. Chao1: the vault's currently-recorded formula
(n−1)/n · f₁²/2f₂ does not match Chao's own 1984 primary, which gives
D + f₁²/(2f₂) with no (n−1)/n term — confirmed by a direct Tier-1 read.
This is an actionable correction for promotion, not a confirmation of the
existing note. Good-Turing f₁/n: still [unverified-quant — needs primary] after this session's attempt; the formula is well-corroborated
across secondary sources (this capture's slide deck, standard NLP/stats
textbook treatments referenced during search, the existing vault note) but
no route to Good's own 1953 text succeeded. Genuine effort was made across
five distinct access routes before stopping.
Further leads
- Chao, A. (1987), "Estimating the population size for capture-recapture data with unequal catchability," Biometrics 43, 783–791 — likely origin of the widely-cited bias-corrected Chao1 form; not read directly this session, only surfaced via search synthesis. Would resolve the
[unverified-mechanism]flag above. - Colwell & Coddington (1994) reportedly coined "Chao2" for the incidence-based sibling estimator — surfaced only in a search snippet, not read.
- Efron & Thisted (1976), "Estimating the number of unseen species: how many words did Shakespeare know?", Biometrika 63, 435–447 — cited inside Chao (1984)'s own reference list; already noted in claim-singletons-are-the-diagnostic-of-the-unseen but worth its own dedicated primary read.
- Harris, B. (1959), "Determining bounds on integrals with applications to cataloging problems," Ann. Math. Statist. 30, 521–548 — the direct mathematical ancestor of Chao (1984); not located or read this session.
- OUP abstract page for Good (1953) is reachable via plain WebFetch (unreceipted) and states "Turing is acknowledged for the most interesting formula in this part of the work" — corroborates but cannot be cited as a quote per the receipts rule until
archive_pagesucceeds against it (currently 403s).
Entity candidates
- Harris, B. (1959) — person/concept — the foundational asymptotic occupancy-bound result Chao's 1984 estimator is explicitly built on ("The method is similar to that taken by Harris (1959)"); the ancestry claim in this capture rests on him, not on any co-author of Chao's (she has none on this paper).
- Alan Turing — person — credited inside Good's 1953 paper (per the OUP abstract fragment and independently per the Weinstein slide's "Why Turing? It was his idea!") as the originator of the unseen-mass formula; already flagged in claim-unseen-species-problem-has-dual-origin-ecology-and-cryptanalysis, reflagged here because this session found a second, independent partial corroboration still short of the primary.
- I. J. Good — person — sole author of the 1953 Biometrika primary; his own text remains unread across this vault after two separate research sessions.
- Anne Chao — person — sole author of the 1984 Scandinavian Journal of Statistics primary, read directly in this session.
- Burnham, K. P. & Overton, W. S. — people/concept — authors of the rival jackknife estimator (1978, 1979) that Chao's 1984 paper explicitly derives as a special case of her own method and benchmarks against in every worked example.
- Bradley Efron — person — originator of the percentile bootstrap method (1981, 1982) that Chao applies to build confidence intervals around her point estimate; adjacent to the vault's existing estimation cluster via Efron–Thisted and James–Stein material.
- Eugene Weinstein — person — presenter of the course slide deck used as the (sub-floor, flagged) stand-in source for Good's formula in this capture.