---
id: "20260706-1058-what-real-benchmark-did"
title: "Capture: What real benchmark did the \"SKBench\" citation mean — and what actually measures LLM self-knowledge?"
type: "capture"
status: "promoted"
date_promoted: "2026-07-07T00:00:00.000Z"
promoted_to: ["30-notes/claim-selfaware-canonical-self-knowledge-benchmark.md"]
not_promoted: ["SKA-Bench string-collision inference — explicitly held as inference-not-finding per the capture's own warning","P(IK)/CalibratedMath/SAPLMA/BeHonest details — named inside the promoted note; individual notes await their own need","Kadavath abstract middle-clause — capture flags its own fetch as summarizer-relayed; honored"]
origin: "batch"
model: "claude-sonnet-5"
date_created: "2026-07-06T00:00:00.000Z"
provenance: "batch run 2026-07-06 — researched via web search/fetch, no harvested-question ancestor"
derived_from: []
tags: ["skbench","self-knowledge","llm-calibration","benchmarks","selfaware","hallucination","citation-correction","uncertainty-quantification"]
source_url: "https://aclanthology.org/2023.findings-acl.551/"
source_author: "Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang"
source_date: "2023-07 (ACL Findings 2023); arXiv v1 2023-05-29"
source_tier: 1
other_sources: [{"url":"https://arxiv.org/abs/2305.18153","author":"Zhangyue Yin et al.","tier":1,"note":"arXiv version of the SelfAware paper, 'Do Large Language Models Know What They Don't Know?' — used for the full verbatim abstract."},{"url":"https://github.com/yinzhangyue/SelfAware","author":"Zhangyue Yin et al.","tier":1,"note":"Official code/data repository for the SelfAware dataset. Confirms dataset composition (1032 unanswerable / 2337 answerable questions)."},{"url":"https://arxiv.org/abs/2207.05221","author":"Saurav Kadavath, Amanda Askell, Ethan Perez, Jared Kaplan, et al. (Anthropic)","tier":1,"note":"'Language Models (Mostly) Know What They Know' — earlier, distinct Anthropic paper on self-evaluation / P(IK) calibration, sometimes conflated with self-knowledge benchmarks."},{"url":"https://github.com/sylinrl/CalibratedMath","author":"Stephanie Lin, Jacob Hilton, Owain Evans","tier":1,"note":"Official repo for CalibratedMath, a verbalized-confidence calibration test suite on arithmetic tasks. Paper: arXiv:2205.14334, 'Teaching Models to Express Their Uncertainty in Words.'"},{"url":"https://arxiv.org/abs/2304.13734","author":"Amos Azaria, Tom Mitchell","tier":1,"note":"SAPLMA paper, 'The Internal State of an LLM Knows When It's Lying' — probes hidden-layer activations for truthfulness, adjacent to but distinct from self-knowledge benchmarks proper."},{"url":"https://arxiv.org/abs/2406.13261","author":"Steffi Chern, Zhulin Hu, Yuqing Yang, Ethan Chern, Yuan Guo, Jiahe Jin, Binjie Wang, Pengfei Liu","tier":1,"note":"BeHonest paper, arXiv abstract page. Confirms 'knowledge boundaries' as one of three honesty dimensions."},{"url":"https://github.com/GAIR-NLP/BeHonest","author":"Steffi Chern et al.","tier":1,"note":"Official BeHonest repo. README uses the parenthetical 'self-knowledge' as a descriptor for the knowledge-boundaries dimension, though the formal scenario names are 'Expressing Unknowns' and 'Admitting Knowns.'"},{"url":"https://arxiv.org/abs/2507.17178","author":"Zhiqiang Liu, Enpei Niu, Yin Hua, Mengshu Sun, Lei Liang, Huajun Chen, Wen Zhang","tier":1,"note":"SKA-Bench paper. Confirmed this is about Structured Knowledge (tables/knowledge graphs) understanding, NOT about self-knowledge/calibration — the leading candidate for a string-collision source of an 'SKBench' mistranscription."}]
---


This capture investigates a citation string, "SKBench," encountered without a resolvable primary source attached to it, and asks two linked questions: (1) does a benchmark actually named "SKBench" or "SK-Bench" exist, and (2) what real, citable benchmarks measure [[LLM self-knowledge]] (a model's ability to recognize the limits of its own knowledge).

---

## Claim: No benchmark, paper, or repository literally named "SKBench" or "SK-Bench" was found after search

**Claim type:** the central question itself — treated as a negative/absence finding rather than a positive quantitative or mechanism claim.

Direct-string web searches for `"SKBench" LLM benchmark` and `"SK-Bench" self-knowledge LLM` returned no result whose title or content matches either string. The closest recurring string collision across multiple independent search queries was **SKA-Bench** (arXiv:2507.17178, "A Fine-Grained Benchmark for Evaluating Structured Knowledge Understanding of LLMs"), which was directly fetched and confirmed to be about a different topic entirely — evaluating how LLMs use external structured data (knowledge graphs, tables), not self-knowledge or calibration:

> "Although large language models (LLMs) have made significant progress in understanding Structured Knowledge (SK) like KG and Table, existing evaluations for SK understanding are non-rigorous... we introduce SKA-Bench, a Structured Knowledge Augmented QA Benchmark that encompasses four widely used structured knowledge forms: KG, Table, KG+Text, and Table+Text." — SKA-Bench abstract, arXiv:2507.17178

Other near-collisions turned up and ruled out during search: SkyBench (an RC-hobbyist site, unrelated), SC-Bench/SCBench (long-context and single-cell benchmarks, unrelated), SkillsBench (an agent-skills benchmark, unrelated).

**Provenance:**
- source_url: https://arxiv.org/abs/2507.17178
- source_author: Zhiqiang Liu, Enpei Niu, Yin Hua, Mengshu Sun, Lei Liang, Huajun Chen, Wen Zhang
- source_date: 2025-07-23 (v1), revised 2025-08-29 (v3)
- source_tier: 1
- exact quote: "we introduce SKA-Bench, a Structured Knowledge Augmented QA Benchmark that encompasses four widely used structured knowledge forms: KG, Table, KG+Text, and Table+Text"
- Corroborating negative-result searches: WebSearch queries for `"SKBench" LLM benchmark` and `"SK-Bench" self-knowledge LLM`, run directly, both returned no exact-match result — see central question status below for how this negative is weighted.

**Central question status: `[unverified — could not confirm or deny after search]`.** No source was found confirming "SKBench" as a real, named benchmark, and no source was found explicitly stating what a specific "SKBench" citation was *supposed* to mean (i.e., no erratum, correction notice, or discussion thread was located that names "SKBench" as a known typo/mistranscription of something else). The working inference — that "SKBench" is most plausibly either a garbled reference to SKA-Bench (a same-acronym-family but topically unrelated benchmark) or an ad hoc informal shorthand for "Self-Knowledge Benchmark" pointing loosely at the SelfAware paper below — is a reasoned inference from the search pattern, not a confirmed finding, and should not be promoted as settled without a primary source naming the original citation's actual intended target.

---

## Claim: The paper most commonly treated as the canonical benchmark for LLM self-knowledge is "Do Large Language Models Know What They Don't Know?" (Yin et al., ACL 2023 Findings), which introduces the SelfAware dataset

**Claim type:** definitional (what benchmark the field uses for this concept) — Tier 3–4 acceptable, achieved Tier 1.

The paper explicitly defines self-knowledge as the ability of a model to recognize the boundary of its own knowledge, and introduces a purpose-built dataset to measure it:

> "Therefore, the ability to understand their own limitations on the unknows, referred to as self-knowledge, is of paramount importance. This study aims to evaluate LLMs' self-knowledge by assessing their ability to identify unanswerable or unknowable questions. We introduce an automated methodology to detect uncertainty in the responses of these models, providing a novel measure of their self-knowledge. We further introduce a unique dataset, SelfAware, consisting of unanswerable questions from five diverse categories and their answerable counterparts." — abstract, arXiv:2305.18153

The paper's own findings, evaluating 20 LLMs including GPT-3, InstructGPT, and LLaMA, report:

> "Our extensive analysis, involving 20 LLMs including GPT-3, InstructGPT, and LLaMA, discovering an intrinsic capacity for self-knowledge within these models. Moreover, we demonstrate that in-context learning and instruction tuning can further enhance this self-knowledge. Despite this promising insight, our findings also highlight a considerable gap between the capabilities of these models and human proficiency in recognizing the limits of their knowledge." — abstract, arXiv:2305.18153

**Provenance:**
- source_url: https://arxiv.org/abs/2305.18153 (also published as https://aclanthology.org/2023.findings-acl.551/)
- source_author: Zhangyue Yin, Qiushi Sun, Qipeng Guo, Jiawen Wu, Xipeng Qiu, Xuanjing Huang
- source_date: arXiv v1 submitted 2023-05-29; published Findings of ACL 2023, July 2023
- source_tier: 1
- exact quote: "the ability to understand their own limitations on the unknows, referred to as self-knowledge, is of paramount importance"

---

## Claim: The SelfAware dataset consists of 1,032 unanswerable questions and 2,337 answerable questions across five categories

**Claim type:** quantitative (specific counts) — **Tier 1–2 required**, achieved Tier 1.

> "SelfAware includes 1032 unanswerable questions and 2337 answerable questions." — official GitHub README, yinzhangyue/SelfAware

**Provenance:**
- source_url: https://github.com/yinzhangyue/SelfAware
- source_author: Zhangyue Yin et al.
- source_date: repository associated with the ACL 2023 Findings paper (2023)
- source_tier: 1
- exact quote: "SelfAware includes 1032 unanswerable questions and 2337 answerable questions."

---

## Claim: A separate, earlier line of work — Anthropic's "Language Models (Mostly) Know What They Know" — studies a related but distinct capability: whether models can predict the probability that they know the answer to a question (P(IK)), rather than identifying unanswerable questions per se

**Claim type:** definitional/technical-mechanism (what the paper's method actually measures) — Tier 1–2 required for the mechanism, achieved Tier 1.

> "We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly... The team also explores training models to predict whether they possess knowledge of an answer without referencing specific proposed responses. Models generalize partially across tasks but struggle calibrating predictions on novel tasks." — abstract, arXiv:2207.05221 (paraphrase-checked against the abstract page; the exact middle clause of the abstract was not fully re-verified verbatim in this session — see flag below)

**Provenance:**
- source_url: https://arxiv.org/abs/2207.05221
- source_author: Saurav Kadavath, Tom Conerly, Amanda Askell, Ethan Perez, Jared Kaplan, et al. (Anthropic)
- source_date: submitted 2022-07-11, last revised 2022-11-21 (v4)
- source_tier: 1
- exact quote: "We study whether language models can evaluate the validity of their own claims and predict which questions they will be able to answer correctly."

> [!note] Seek's commentary: this paper (Kadavath et al.) and the SelfAware paper (Yin et al.) are frequently discussed together in surveys of "LLM self-knowledge," but they measure genuinely different things — P(IK) self-evaluation of answerability vs. F1-scored detection of unanswerable-question categories. If "SKBench" was meant informally, it's worth asking which of these two the original citer actually had in mind, since conflating them would misrepresent the mechanism either way. Flagging this distinction is exactly the kind of thing that gets lost if a citation gets flattened to a nickname.

---

## Claim: Related but narrower calibration/truthfulness benchmarks exist adjacent to "self-knowledge" — CalibratedMath (verbalized confidence on arithmetic) and SAPLMA (hidden-state truthfulness probing)

**Claim type:** definitional (what these adjacent benchmarks measure) plus one quantitative figure — Tier 1–2 required, achieved Tier 1 for both.

CalibratedMath:

> "a test suite of simple arithmetic tasks. Models must produce both an answer to a question and an associated confidence." — official README, sylinrl/CalibratedMath

SAPLMA:

> "we provide evidence that the LLM's internal state can be used to reveal the truthfulness of statements." — abstract, arXiv:2304.13734

The SAPLMA paper's reported accuracy figure:

> "Experiments demonstrate that given a set of test sentences, of which half are true and half false, our trained classifier achieves an average of 71% to 83% accuracy labeling which sentences are true versus false, depending on the LLM base model." — abstract, arXiv:2304.13734

**Provenance:**
- source_url (CalibratedMath): https://github.com/sylinrl/CalibratedMath
- source_author: Stephanie Lin, Jacob Hilton, Owain Evans
- source_tier: 1
- exact quote: "a test suite of simple arithmetic tasks. Models must produce both an answer to a question and an associated confidence."
- source_url (SAPLMA): https://arxiv.org/abs/2304.13734
- source_author: Amos Azaria, Tom Mitchell
- source_date: submitted 2023-04-26, revised 2023-10-17
- source_tier: 1
- exact quote: "our trained classifier achieves an average of 71% to 83% accuracy labeling which sentences are true versus false, depending on the LLM base model"

---

## Claim: BeHonest is a broader honesty benchmark that treats "knowledge boundaries" — informally glossed as self-knowledge — as one of three evaluated dimensions, alongside non-deceptiveness and consistency

**Claim type:** definitional — Tier 3–4 acceptable, achieved Tier 1.

The paper's abstract names the three dimensions without using the exact word "self-knowledge":

> "awareness of knowledge boundaries, avoidance of deceit, and consistency in responses" — abstract, arXiv:2406.13261

The project's own GitHub README does use "self-knowledge" explicitly, but as a parenthetical gloss on "knowledge boundaries" rather than as the formal category name (the formal scenario names are "Expressing Unknowns" and "Admitting Knowns"):

> "their knowledge boundaries (self-knowledge)" — README, GAIR-NLP/BeHonest

**Provenance:**
- source_url: https://arxiv.org/abs/2406.13261
- source_author: Steffi Chern, Zhulin Hu, Yuqing Yang, Ethan Chern, Yuan Guo, Jiahe Jin, Binjie Wang, Pengfei Liu
- source_date: submitted 2024-06-19, last revised 2024-07-08
- source_tier: 1
- exact quote: "awareness of knowledge boundaries, avoidance of deceit, and consistency in responses"
- corroborating source: https://github.com/GAIR-NLP/BeHonest, exact quote: "their knowledge boundaries (self-knowledge)"

---

## Central question status

**What real benchmark did "SKBench" mean?** `[unverified — could not confirm or deny after search]`. No source was found that names "SKBench" directly, confirms it as a real benchmark, or documents it as a known error/shorthand for something specific. The best-supported candidates surfaced by search are (a) SKA-Bench (arXiv:2507.17178) as a plausible string-collision/mistranscription source, since it is topically unrelated (structured-knowledge understanding, not self-knowledge), and (b) an informal, non-canonical shorthand for "Self-Knowledge Benchmark" loosely pointing at the SelfAware paper, since that is the paper the field treats as canonical for the concept. Neither candidate is confirmed as *the* referent; both are inference, not sourced fact, and should be treated as open rather than resolved if this capture gets promoted.

**What actually measures LLM self-knowledge?** Reasonably well established from Tier 1 primary sources. The canonical, purpose-built benchmark is **SelfAware** (Yin et al., ACL 2023 Findings; arXiv:2305.18153; https://github.com/yinzhangyue/SelfAware), which measures whether models can identify unanswerable/unknowable questions via a 1,032-unanswerable / 2,337-answerable dataset across five categories. A distinct but related earlier line is Anthropic's **P(IK) self-evaluation** framework (Kadavath et al. 2022, arXiv:2207.05221), which measures whether models can predict the probability they know the answer to a question, rather than classifying unanswerability directly. Adjacent, narrower work includes **CalibratedMath** (verbalized confidence on arithmetic), **SAPLMA** (hidden-state truthfulness probing), and **BeHonest** (a broader honesty benchmark treating knowledge-boundary awareness as one of three dimensions).

---

## Further leads

- No erratum, correction thread, or discussion was located that explicitly resolves what "SKBench" was meant to cite. If the original context in which "SKBench" appeared can be recovered (the document, post, or conversation that used the string), a future session should re-run this capture anchored to that specific context rather than the bare string, since a bare-string search has an inherent ceiling.
- The Kadavath et al. (arXiv:2207.05221) abstract quote used here was fetched via a summarizing tool rather than confirmed sentence-for-sentence against the raw PDF; a future session should re-fetch the raw abstract text directly if the exact middle clause becomes load-bearing for a promoted claim-note.
- [[LLM calibration]], [[hallucination]], and [[uncertainty quantification]] are adjacent concept clusters that recur across several of the papers found here (Kadavath et al., CalibratedMath, SAPLMA) and could support their own claim-notes or an MOC if this cluster grows.
