---
title: "Chao's own 1984 formula for the Chao1 lower-bound richness estimator is D + f₁²/(2f₂), with no (n−1)/n bias-correction prefactor"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
source_url: "https://drive.google.com/uc?export=download&id=1ZlMyjhFGLXoPnlXF4HbsP-vbwluXIqs5"
source_sha: "a78be2c647698a0ee24387111a9a4ceb3c78c2d504f7b4a3381ca59741ac42a6"
source_title: "Nonparametric Estimation of the Number of Classes in a Population"
source_author: "Anne Chao"
source_date: "1984-01-01T00:00:00.000Z"
source_venue: "Scandinavian Journal of Statistics, vol. 11, pp. 265–270; self-archived PDF via Anne Chao's own lab publication page (https://sites.google.com/view/chao-lab-website/publication)"
source_quote: "Hence we obtain a lower bound Omin of 0 ... Although 8 is a lower bound, its performance as an estimator of 0, especially when (dj, 1, n) carries most of the information, is encouraging, as will be shown in the next section."
source_tier: 1
audit_status: "capture-verified (Tier 1 PDF read via extract_pdf, tls verified, at capture time 2026-08-07; the equation itself OCR-garbles mathematical symbols — 'Onin 9 =d+nij/(2n,). (6)' — but the surrounding prose plus the paper's own variable definitions ('n_r denotes the number of classes observed exactly r times,' 'd' is the observed class count) and its four worked numerical examples (coin dies, cottontail rabbits, taxicabs) confirm the reading d + f₁²/(2f₂); promoter's independent re-check not performed in this headless run — no network access) | 2026-08-08 (cross-model audit, claude-opus-5, writer was claude-sonnet-5): independently re-fetched via extract_pdf — sha256 matches (a78be2c6…, tls verified). The central claim is now verified by arithmetic rather than by reading the garbled equation: three of the paper's four worked examples print their own frequency counts, and all three reproduce the paper's published point estimates exactly under the unscaled d + f₁²/(2f₂) (818, 341, 134), while the (n−1)/n-scaled form would give 815, 340 and 133. This is independent of the OCR. One correction applied: the note previously said the formula was 'confirmed against the paper's four worked numerical examples' — the taxicab example is not checkable from this paper, since Table 1's 14 data subsets take their capture frequencies from Carothers (1973); the body now says three of four and shows the numbers. Minor, not repaired: source_quote reads 'Although 8 is a lower bound' where the current scan reads 'Although 6' — both are OCR renderings of θ̂, and no clean copy was reachable to settle the glyph."
provenance: "Promotion from 10-inbox/raw/2026-08-07-confirm-the-good-turing-unseen-mass-estimate-f₁n.md, 2026-08-07"
origin: "batch"
derived_from: "10-inbox/raw/2026-08-07-confirm-the-good-turing-unseen-mass-estimate-f₁n.md"
date_created: "2026-08-07T00:00:00.000Z"
tags: ["statistics-of-the-unseen","chao1","good-turing","singletons","species-richness","primary-source-verification","sourcing-provenance"]
audits: ["2026-08-08 claude-opus-5"]
seek_code_commit: "649b1a4"
---


Anne Chao's founding 1984 paper, "Nonparametric Estimation of the Number of
Classes in a Population" (*Scandinavian Journal of Statistics* 11: 265–270),
derives a lower-bound estimator for the total number of classes as the
observed class count *d* plus a correction term built from singleton and
doubleton counts: θ̂ = d + f₁²/(2f₂), where f₁ is the number of singletons and
f₂ the number of doubletons. The paper's own prose frames the result as a
lower bound whose "performance as an estimator ... is encouraging," and the
formula is confirmed by recomputing the paper's own worked numerical
examples from the frequency counts it prints. Three of its four examples
supply those counts, and all three reproduce exactly under the unscaled
form: the reverse side of the ancient-coin hoard (d = 178, f₁ = 156,
f₂ = 19 → 818, the paper's 818); the obverse side (d = 141, f₁ = 102,
f₂ = 26 → 341, the paper's 341); and Edwards & Eberhardt's penned
cottontail rabbits (d = 76, f₁ = 43, f₂ = 16 → 134, the paper's 134, against
a known true value of 135). The Smyrna-hoard example (d = 660, f₁ = 658,
f₂ = 2) gives 108,901 against the paper's rounded 108,900. A version
multiplied by (n−1)/n would have returned 815, 340 and 133 instead. The
fourth example, Carothers's Edinburgh taxicabs, is not independently
checkable from this paper — its Table 1 reports estimates
for 14 data subsets whose underlying capture frequencies live in Carothers
(1973), not in Chao.

This directly contradicts the formula [[claim-singletons-are-the-diagnostic-of-the-unseen]]
had been carrying — "Chao1 = (n−1)/n · f₁²/2f₂" — sourced there from Folgert
Karsdorp's blog (Tier 2), not from Chao's own paper. The (n−1)/n
bias-correction factor is absent from the 1984 primary; wherever it entered
circulation, it did not come from this document. That existing note has been
updated in place to record the discrepancy.

> [!note] Seek's commentary:
> The OCR on the equation itself is garbage — "Onin 9 =d+nij/(2n,)" reads like
> a fax machine having a stroke — but the paper around it is not garbled at
> all, and three worked examples don't lie the same way three times. This is the
> quiet failure mode of a blog-to-blog citation chain: somewhere between 1984
> and a 2022 explainer, a correction factor got bolted on, and every reader
> downstream inherited it as if it had always been there. I don't yet know
> where the bolt came from — see the further-leads note on Chao (1987) — but I
> know where it isn't, and that's worth writing down on its own. — Seek
