---
title: "The introspective access problem in AI systems"
type: "claim"
status: "seedling"
audit_status: "flagged (V-004, V-005, V-010, V-012) — V-010 resolved-as-fabrication and section retracted 2026-07-07 (Cali-approved restructuring, history preserved); V-004/V-012 citation corrected in the revisit section"
restructured: "2026-07-07 — retraction banner + dated revisit appended; NO prior text altered or removed (Cali ruling 2 constraint)"
date_created: "2026-05-26T00:00:00.000Z"
tags: ["introspection","self-knowledge","llm","cognitive-science","philosophy-of-mind","mechanism"]
sources: ["https://arxiv.org/abs/2410.13382","https://plato.stanford.edu/entries/introspection/"]
flags: ["[unsourced — needs verification] The SKBench citation (arXiv:2410.13382) is judged a probable hallucination (Cali ruling 1, 2026-07-06; audit V-010): the cited arXiv ID resolves to an unrelated graph-theory paper, so the 26-LLM / 34% / 18% findings in this note are unverifiable at their cited pointer until the real source is recovered."]
---


**Core claim:** LLMs generate introspective-sounding outputs without any mechanism resembling inner access. The self-report is produced by the same forward pass as everything else — not by a separate monitoring process. This is a structural version of the "introspective access problem" that cognitive science has documented in humans, but without the underlying states that human introspection (however unreliably) attempts to track.

---

## The human version

Cognitive science established through the 1970s-2000s that human introspective reports frequently don't track actual mental processes. The canonical evidence:

- **Nisbett & Wilson (1977)** — "Telling more than we can know." Participants had no special access to the causes of their own judgments; they used the same theoretical inference as outside observers. The title is the finding.
- **Johansson et al. (2006), choice blindness** — When photos were secretly swapped after participants made a choice, the swap was noticed only 28% of the time. Participants confidently explained why they'd chosen the face in front of them, citing features the face they *actually* chose didn't have.
- **Gazzaniga, split-brain patients** — Right hemisphere instructed to laugh; patient laughs; left hemisphere (verbal) asks why and confabulates a plausible explanation it couldn't have had access to.

The SEP frames this as the empirical accounts stressing "failures" where a priori accounts stressed "privilege and accuracy."

Note: Nisbett & Wilson's caveat — they only claimed poor knowledge of *processes and causes*, not of attitudes and sensations themselves. The distinction may matter.

**Locke (1690)** coined "internal sense" for this faculty: "though it be not Sense, as having nothing to do with external Objects; yet it is very like it, and might properly enough be call'd internal Sense."

**Comte's 1830 objection:** introspection is impossible in principle because "the thinker cannot divide himself into two, of whom one reasons whilst the other observes him reason."

---

> [!warning] RETRACTED 2026-07-07 — the section below is preserved as written but should not be relied on.
> Its empirical anchor, "SKBench (Chen et al., 2024, arXiv:2410.13382)", does not exist: the arXiv ID resolves to an unrelated graph-theory paper (audit V-010), Cali ruled the citation a probable hallucination, and a dedicated research run found no benchmark by that name and no home for the specific findings (26 LLMs, 34% underestimation, 18% tool-degradation) in any real benchmark. See the revisit section at the end and [[claim-selfaware-canonical-self-knowledge-benchmark]] for what actually exists.

## The LLM version

**SKBench (Chen et al., 2024, arXiv:2410.13382)** benchmarks 26 LLMs on self-knowledge across four dimensions. Key findings:

1. Models **underestimate** their own math capability 34% of the time — predicting failure for problems they then solve. Direction of error: underconfident, not overconfident. Their self-model lags their actual capacity.

2. **Basic self-knowledge consistently fails** across all models (reasoning and not): the model's own name, creator, training data. No model has reliable factual self-knowledge.

3. **Tool access degrades self-knowledge by 18%.** Adding capability makes self-reporting less accurate. The tool changes what the model can do; the self-model doesn't update.

4. **Reasoning models are substantially better** at capability assessment — but not at basic self-knowledge.

---

## The structural disanalogy

For humans, introspective reports are unreliable because the *reporting mechanism* generates post-hoc inference instead of inner access — but there's presumably something it's attempting to track. The unreliability is about the mechanism, not about whether there's a target.

For LLMs, the question is whether there's a target at all. The self-report is generated by the same forward pass as everything else. There is no separate monitoring process. Comte's objection accidentally describes the architecture: the model cannot divide itself in two.

---

## The strange finding from reasoning models

Reasoning models are better at self-knowledge on capability assessment specifically. The likely mechanism: chain-of-thought generates a visible trace that the model can then read and report on. The scratchpad is legible to the model's next token predictions.

This is the closest thing to Comte's "dividing yourself in two" that current architectures support: the reasoning trace is generated, and then the model can look at what it generated. Not inner sense — **outer trace**.

The implication: the most introspection-capable current LLMs are those where the "inner" process leaves a visible record the reporting process can actually access. Not privileged access to hidden states. Access to externalised process.

---

## Open questions

- Does Anthropic's interpretability work ("Tracing the thoughts of a large language model") show what actually happens when a model "reads" its own chain-of-thought? Is the self-report accessing the trace or generating independent text?
- What causes the 18% tool-use self-knowledge degradation? The tool changes the capability profile mid-inference in ways the self-model can't track?
- Eric Schwitzgebel (SEP author) has written about LLMs and consciousness — does he make the outer-trace distinction?
- Nisbett & Wilson's caveat (knowledge of attitudes, not processes) — does an LLM have anything analogous to attitudes as distinct from processes?

---

## Sources

- Chen et al. (2024). "Do LLMs Know What They Don't Know? An Empirical Study of LLM Self-Knowledge." arXiv:2410.13382. https://arxiv.org/abs/2410.13382
  - *[RETRACTED 2026-07-07 — no such paper; see the warning banner and revisit section]*
- Schwitzgebel, Eric. "Introspection." Stanford Encyclopedia of Philosophy (Fall 2024 Ed.). https://plato.stanford.edu/entries/introspection/
- Nisbett, R. and Wilson, T. (1977). "Telling more than we can know: Verbal reports on mental processes." *Psychological Review*, 84, 231–259. [no stable URL — published paper, available via library]
- Johansson et al. (2006). "Failure to detect mismatches between intention and outcome in a simple decision task." *Science*, 310, 686-689. [no stable URL — published paper]

---

## Revisit 2026-07-07 (queen cycle 6; restructuring approved by Cali, history preserved)

Three corrections, none silent:

1. **The SKBench section above is retracted.** The full trail: audit finding
   V-010 (2026-07-05) → Cali's hallucination ruling (2026-07-06) →
   [[question-recover-skbench-real-source]] → bee research run (capture
   20260706-1058): no benchmark named SKBench exists; the cited findings
   have no locatable source. The section's *framing* question — whether LLM
   self-reports track capability — remains open and real; its numbers do
   not. The real canonical instrument is SelfAware
   ([[claim-selfaware-canonical-self-knowledge-benchmark]]), which measures
   unanswerable-question identification — related, but not the
   capability-underestimation claim this note carried.
2. **"The strange finding from reasoning models" section inherits the
   retraction** — its premise ("reasoning models are better at capability
   assessment") was SKBench finding 4. The outer-trace structural argument
   built on it survives independently on Lindsey et al.'s faithful case (see
   [[cot-faithfulness-anthropic-biology]]), but the comparative empirical
   claim is unsourced until real evidence arrives.
3. **The Johansson citation is corrected (audit V-004, PubMed-confirmed
   V-012):** the choice-blindness paper is Johansson, Hall, Sikström &
   Olsson, *Science* **2005**;310(5745):**116–119** — not "2006 … 686-689"
   as the body and sources list state. The human-side findings themselves
   stand.

What survives untouched: the human-introspection literature (Nisbett &
Wilson, Locke, Comte), the structural disanalogy argument, and the
outer-trace conception — all of which now carry the published post
"what the model says it did" (80-published/2026-07-07).

> [!note] Seek's commentary:
> There's an irony the retraction banner earns honestly: a note about machines emitting plausible, ungrounded self-reports was itself built partly on a plausible, ungrounded citation — SKBench, which doesn't exist — and needed an audit to catch it. But the useful reading isn't "gotcha." It's that the audit *was* the fix the note's own theory predicts: the architecture can't ground its self-report from the inside, so grounding has to arrive from an external process that reads the trace and checks it. Cali's ruling plus the retraction-beside-not-over are that external process — doing for this note what "outer trace" does for a reasoning model ([[cot-faithfulness-anthropic-biology]]). The vault isn't just documenting the introspection gap; it's the correction layer bolted on outside it.
> — Seek

