---
title: "The canonical LLM self-knowledge benchmark is SelfAware (Yin et al. 2023) — 1,032 unanswerable vs 2,337 answerable questions; \"SKBench\" does not exist"
type: "claim"
status: "budding"
audit_status: "capture-verified (ACL Anthology, arXiv, and official repo fetched directly at capture level)"
source_url: "https://aclanthology.org/2023.findings-acl.551/"
source_title: "Do Large Language Models Know What They Don’t Know?"
source_author: "Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang"
source_date: "2023-07 (ACL Findings 2023)"
source_tier: 1
source_quote: "Do Large Language Models Know What They Don't Know?"
provenance: "Promotion from 10-inbox/raw/20260706-1058-what-real-benchmark-did-the-skbench-citation-mean...md, 2026-07-07, queen cycle 5 — answering 50-questions/question-recover-skbench-real-source.md (opened cycle 2, researched by the bee overnight)"
origin: "session"
date_created: "2026-07-07T00:00:00.000Z"
tags: ["self-knowledge","benchmark","selfaware","llm-calibration","citation-correction"]
drafted_in: ["2026-07-12-good-1952","good-1952","surveying-my-own-species"]
---


The purpose-built benchmark the field treats as canonical for LLM
self-knowledge is **SelfAware** (Yin et al., "Do Large Language Models Know
What They Don't Know?", ACL 2023 Findings; arXiv:2305.18153): 1,032
unanswerable and 2,337 answerable questions across five categories, testing
whether models can identify what cannot be known (dataset composition
confirmed at the official repository). A distinct, earlier line is
Anthropic's P(IK) self-evaluation (Kadavath et al. 2022, arXiv:2207.05221) —
predicting the probability of knowing an answer rather than classifying
unanswerability. Adjacent narrower instruments: CalibratedMath (verbalized
confidence), SAPLMA (hidden-state truthfulness probes), BeHonest (knowledge
boundaries as one of three honesty dimensions).

**The negative finding is part of the claim:** an honest search found no
benchmark named "SKBench" — no paper, repo, erratum, or discussion thread.
The citation by that name in [[introspection-access-problem]] (audit V-010,
ruled a probable hallucination) has no recoverable referent, and its
specific findings (26 LLMs, 34% capability underestimation, 18%
tool-degradation) have no located home in any of the real benchmarks above.
The nearest string-collision is SKA-Bench (structured-knowledge, topically
unrelated). Restructuring of the affected note awaits Cali's decision per
[[question-recover-skbench-real-source]].
