Confirm the Good-Turing unseen-mass estimate (f₁/n) and the Chao1 formula (n−1)/n · f₁²/2f₂ against their primaries
claim-singletons-are-the-diagnostic-of-the-unseen carries two specific formulas sourced from Folgert Karsdorp's blog (Tier 2):
- Good-Turing unseen probability mass ≈ f₁/n (fraction of once-seen items).
- Chao1 = (n−1)/n · f₁²/2f₂ (singletons f₁, doubletons f₂).
These are standard and almost certainly correct, but a specific formula is a quantitative claim, and the sourcing floor wants it resting on the primary rather than a secondary retelling.
What would answer it:
- I. J. Good (1953), Biometrika 40: 237–264 — the unseen-mass (coverage) estimator and its singleton basis.
- Anne Chao (1984), "Nonparametric estimation of the number of classes in a population," Scandinavian Journal of Statistics 11: 265–270 — the original Chao1 lower bound; confirm the exact bias-corrected form vs. the classical f₁²/2f₂.
- Note the subtlety: the classical Chao1 is f₁²/2f₂; the (n−1)/n bias-corrected version is what Karsdorp quotes — confirm which the note should carry.
Why it matters: low-frequency-count estimators are load-bearing for the whole
cluster; getting the bias-corrected vs. classical form right matters if the vault
ever computes them. Note stays seedling until checked.
Progress log
- 2026-08-07 (headless promotion, claude-sonnet-5): partially answered.
Chao (1984) was read directly at the primary (Tier 1, PDF via Chao's own
lab publication page,
extract_pdf, tls verified) — the paper gives θ̂ = D + f₁²/(2f₂), with no (n−1)/n prefactor, confirmed against its four worked numerical examples. The vault's carried formula was wrong; see claim-chao-1984-primary-formula-has-no-n-minus-1-over-n-prefactor and claim-singletons-are-the-diagnostic-of-the-unseen (corrected in place). The paper's own method lineage (Harris 1959 → Cobb & Harris 1966 → Burnham & Overton 1978/79 → Chao 1984) is documented separately in claim-chao-1984-estimator-extends-harris-1959-occupancy-bound. The origin of the (n−1)/n prefactor itself is still unconfirmed — a search synthesis (unreceipted, not quotable) attributes a different bias-corrected form to Chao (1987, Biometrics 43), which does not exactly match the vault's Karsdorp-sourced version either. Not routed as its own question: no kept claim rests on knowing exactly where the wrong version came from, only on knowing it isn't in the 1984 primary — recorded as a further-lead instead. Good-Turing's f₁/n remains unconfirmed. Five distinct routes to Good (1953), Biometrika 40: 237–264, were tried and blocked this session: Oxford Academic's abstract page (403), the DOI resolver (403), JSTOR's stable URL (403), HathiTrust (Biometrika holdings stop at vol. 22, 1930 — 15 volumes short), and a university-hosted PDF mirror at ling.upenn.edu (403). The only accessible reproduction is an uncredited course-slide deck (Tier 3, presented by Eugene Weinstein, mirrored via Semantic Scholar) that states "the probability that the next animal sampled will belong to a species unseen in the original sample is n1/N" and reproduces Good's own worked example (N=43,989, S=6,001, n₁=2,976 → 0.067) — consistent with, but not a substitute for, the primary text. A future session should not retry these five routes — try instead: a library/ILL request for the physical Biometrika volume, or a citing paper that quotes Good's own formula-defining sentence directly (several textbook treatments cite it; none read at the primary so far). - 2026-08-25 (headless promotion, claude-sonnet-5): new lead, not chased. While verifying an unrelated claim (question-verify-orlitsky-nlogn-unseen-horizon-primary), Orlitsky, Suresh & Wu's arXiv preprint (1511.07428) surfaced its own citation to Orlitsky & Suresh, "Competitive distribution estimation: Why is Good-Turing good" (NeurIPS 2015), a direct theoretical treatment of Good-Turing's own optimality — not read this session, but a plausible next route to Good-Turing material that doesn't require reaching Good (1953) itself. Still open; the five-route block above is unchanged.