---
title: "The count of things seen exactly once — ecology's 'singletons', linguistics' 'hapax legomena' — is the diagnostic of the unseen: Good-Turing sets the unseen mass at roughly f₁/n"
type: "claim"
status: "seedling"
audit_status: "flagged (unverified-quant — the Good-Turing unseen-mass estimate f₁/n and the Chao1 richness formula (n−1)/n · f₁²/2f₂ are carried here from Karsdorp's blog (Tier 2). These are standard textbook formulas, but a specific formula is a quantitative claim whose primary homes are Good 1953 (Biometrika) and Chao 1984 (Scand. J. Statistics), not a blog. Routed to [[question-verify-good-turing-chao1-formulas-primary]])"
writer_model: "claude-opus-4-8"
source_url: "https://www.karsdorp.io/posts/20220309103709-good_turing_as_an_unseen_species_model/"
source_title: "Demystifying Chao1 with Good-Turing"
source_author: "Folgert Karsdorp, 'Good-Turing as an unseen-species model'"
source_date: "2022-03-09T00:00:00.000Z"
source_quote: "Chao1 = (n−1)/n · f₁²/2f₂"
source_tier: 2
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md, 2026-07-12 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["statistics-of-the-unseen","good-turing","singletons","hapax-legomena","chao1","cross-domain-bridge"]
---


Across the fields that share the unseen-species problem
([[claim-unseen-species-problem-has-dual-origin-ecology-and-cryptanalysis]]), the
quantity that tells you how much you have *not* seen is the number of things you
have seen **exactly once**. Ecology calls these **singletons**; linguistics calls
them **hapax legomena** — words appearing once in a corpus. The same statistic,
two field-names, is the observable proxy for the invisible tail.

Two estimators formalize it. Good-Turing sets the total probability mass of
never-seen items at approximately **f₁/n** — the fraction of the sample made of
once-seen items — the intuition being that if many things have shown up only once,
many more are still waiting to show up at all. The **Chao1** lower-bound estimator
of total richness uses singletons *and* doubletons: Chao1 = (n−1)/n · f₁²/2f₂,
where f₁ is the count of singletons and f₂ of doubletons. When there are no
doubletons the community is well-sampled; a large singleton-to-doubleton ratio
signals a large hidden tail.

This diagnostic is what powers the sample-coverage machinery in
[[claim-paleobiology-reinvented-coverage-based-rarefaction-as-quorum-subsampling]]:
coverage is estimated from exactly these low-frequency counts. It is also the hook
that connects the problem to authorship attribution — Efron and Thisted used the
same once-seen statistics to estimate the words Shakespeare knew but never wrote —
which sits near the vault's estimation cluster
([[claim-james-stein-estimator-uniformly-dominates-the-sample-mean]],
[[claim-efron-baseball-shrinkage-halved-batting-average-prediction-error]]).

**`[unverified-quant — needs primary]`.** The two formulas are standard but are
carried here from a Tier-2 blog rather than their primaries (Good 1953; Chao
1984); verification is routed to
[[question-verify-good-turing-chao1-formulas-primary]]. The note stays `seedling`.
