Verify the n·log n unseen-prediction horizon and its 'best possible' optimality against Orlitsky, Suresh & Wu (PNAS 2016) and Valiant & Valiant
claim-unseen-mass-is-predictable-only-to-n-log-n attributes to Orlitsky, Suresh & Wu (PNAS 2016) the result that from n samples the unseen is predictable only about n·log n observations further, and that this range "is the best possible." The PNAS paper is Tier 1, but it was not read in the headless promotion — the bound and the quoted phrase came through the hop capture. Valiant & Valiant's concurrent proof is cited secondhand.
What would answer it:
- Orlitsky, Suresh & Wu (2016), "Optimal prediction of the number of unseen species," PNAS 113(47): 13283–13288 — confirm the n·log n statement, the exact optimality wording, and what "predict" means precisely (which loss, which regime).
- Valiant & Valiant (2016), "Estimating the unseen: improved estimators for entropy and other properties" / the "bird in the hand is worth log n in the bush" result — confirm it is genuinely a concurrent, independent proof of the same limit.
- Pin down whether n·log n is the horizon for prediction of new species specifically vs. estimation of other distribution properties.
Why it matters: this bound is the load-bearing caveat the note adds to the seed's
No-Change Assumption (claim-kb-completeness-toolkit-cardinality-nca-recall); a
quantitative limit used to qualify a completeness heuristic must rest on its primary.
Note stays seedling until read.