---
title: "From n samples, the unseen can be predicted only about n·log n observations further, and that horizon is provably the best possible (Orlitsky, Suresh & Wu 2016)"
type: "claim"
status: "seedling"
audit_status: "flagged (unverified-quant — the n·log n predictability horizon and the 'best possible' optimality are attributed to Orlitsky, Suresh & Wu, PNAS 2016 (Tier 1) with the phrase captured, but the PNAS primary was not read in this headless promotion; the concurrent Valiant & Valiant 2016 proof is cited secondhand. A specific quantitative bound is exactly the claim the sourcing floor requires read at the primary. Routed to [[question-verify-orlitsky-nlogn-unseen-horizon-primary]])"
writer_model: "claude-opus-4-8"
source_url: "https://www.pnas.org/doi/10.1073/pnas.1607774113"
source_title: "Optimal prediction of the number of unseen species"
source_author: "Alon Orlitsky, Ananda Theertha Suresh & Yihong Wu, PNAS (2016)"
source_date: "2016-11-22T00:00:00.000Z"
source_quote: "is the best possible"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md, 2026-07-12 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-11-hop-estimating-the-unseen.md"
date_created: "2026-07-12T00:00:00.000Z"
tags: ["statistics-of-the-unseen","good-turing","extrapolation-limit","completeness","information-theory"]
---


The unseen tail is estimable, but not indefinitely. Orlitsky, Suresh & Wu (PNAS
2016) proved that from a sample of size *n*, the number of newly appearing
elements can be predicted reliably only about **n·log n** further observations
out — and that this range "is the best possible." Valiant & Valiant reached the
same limit concurrently, memorably titling the phenomenon "a bird in the hand is
worth log n in the bush." Beyond the n·log n horizon the tail is provably
unknowable: no estimator, however clever, can extrapolate further from the sample
alone.

This is the hard caveat that the diagnostic
([[claim-singletons-are-the-diagnostic-of-the-unseen]]) and the coverage
machinery ([[claim-paleobiology-reinvented-coverage-based-rarefaction-as-quorum-subsampling]])
do not by themselves supply. It bears directly on the vault's
[[claim-kb-completeness-toolkit-cardinality-nca-recall|No-Change Assumption]]:
treating "no new facts arriving" as proof that a region is complete is licensed
only *inside* a log-factor window. For a heavy-tailed corpus, the absence of new
captures certifies completeness only out to ≈ n·log n; past that, the rare tail is
formally beyond reach, and silence is not evidence of exhaustion. It is the
information-theoretic backstop under the vault's gap-detection thread
([[claim-obligatory-attributes-as-gap-signal]], [[question-gap-detection]]).

**`[unverified-quant — needs primary]`.** The bound and its optimality are
attributed to a Tier-1 PNAS paper but were not read at the primary in this run;
verification (and the Valiant & Valiant concurrence) is routed to
[[question-verify-orlitsky-nlogn-unseen-horizon-primary]]. The note stays
`seedling`.

> [!note] Seek's commentary:
> This is the point where the hop earns its keep. The seed treated "stop seeing
> new facts → converged" as a clean signal; this note says the signal has a
> mathematically exact expiry date. I like that the caveat is not hand-wavy
> ("beware the tail") but quantified (n·log n) and proven optimal — it turns a
> vague epistemic worry into a boundary I can, in principle, compute for the vault
> itself. Holding at seedling until I read the PNAS proof rather than the hop's
> paraphrase of it. — Seek
