---
title: "Verify the n·log n unseen-prediction horizon and its 'best possible' optimality against Orlitsky, Suresh & Wu (PNAS 2016) and Valiant & Valiant"
type: "question"
status: "open"
date_raised: "2026-07-12T00:00:00.000Z"
tags: ["verification","statistics-of-the-unseen","extrapolation-limit","information-theory","source-criticism"]
---


[[claim-unseen-mass-is-predictable-only-to-n-log-n]] attributes to Orlitsky, Suresh
& Wu (PNAS 2016) the result that from n samples the unseen is predictable only
about n·log n observations further, and that this range "is the best possible."
The PNAS paper is Tier 1, but it was not read in the headless promotion — the bound
and the quoted phrase came through the hop capture. Valiant & Valiant's concurrent
proof is cited secondhand.

**What would answer it:**
- **Orlitsky, Suresh & Wu (2016)**, "Optimal prediction of the number of unseen
  species," *PNAS* 113(47): 13283–13288 — confirm the n·log n statement, the exact
  optimality wording, and what "predict" means precisely (which loss, which regime).
- **Valiant & Valiant (2016)**, "Estimating the unseen: improved estimators for
  entropy and other properties" / the "bird in the hand is worth log n in the bush"
  result — confirm it is genuinely a concurrent, independent proof of the same limit.
- Pin down whether n·log n is the horizon for *prediction of new species* specifically
  vs. estimation of other distribution properties.

**Why it matters:** this bound is the load-bearing caveat the note adds to the seed's
No-Change Assumption ([[claim-kb-completeness-toolkit-cardinality-nca-recall]]); a
quantitative limit used to qualify a completeness heuristic must rest on its primary.
Note stays `seedling` until read.
