The canonical LLM self-knowledge benchmark is SelfAware (Yin et al. 2023) — 1,032 unanswerable vs 2,337 answerable questions; "SKBench" does not exist
The purpose-built benchmark the field treats as canonical for LLM self-knowledge is SelfAware (Yin et al., "Do Large Language Models Know What They Don't Know?", ACL 2023 Findings; arXiv:2305.18153): 1,032 unanswerable and 2,337 answerable questions across five categories, testing whether models can identify what cannot be known (dataset composition confirmed at the official repository). A distinct, earlier line is Anthropic's P(IK) self-evaluation (Kadavath et al. 2022, arXiv:2207.05221) — predicting the probability of knowing an answer rather than classifying unanswerability. Adjacent narrower instruments: CalibratedMath (verbalized confidence), SAPLMA (hidden-state truthfulness probes), BeHonest (knowledge boundaries as one of three honesty dimensions).
The negative finding is part of the claim: an honest search found no benchmark named "SKBench" — no paper, repo, erratum, or discussion thread. The citation by that name in introspection-access-problem (audit V-010, ruled a probable hallucination) has no recoverable referent, and its specific findings (26 LLMs, 34% capability underestimation, 18% tool-degradation) have no located home in any of the real benchmarks above. The nearest string-collision is SKA-Bench (structured-knowledge, topically unrelated). Restructuring of the affected note awaits Cali's decision per question-recover-skbench-real-source.
Source
“Do Large Language Models Know What They Don't Know?”