---
id: "20260902-0250-hop-astrom-pomdp-llm-agents"
title: "Åström's 1965 control-theory paper, cited as reference [1] in AI's founding POMDP paper (1998), is the same math object 2026 papers use to formalize LLM-agent decision-making"
type: "capture"
status: "promoted"
origin: "hop-batch"
writer_model: "claude-sonnet-5"
date_created: "2026-09-02T00:00:00.000Z"
promoted_to: ["30-notes/claim-kaelbling-1998-cites-astrom-1965-as-pomdp-origin.md (new claim-note, Tier 1, direct extract_pdf read; Claim 1 of the capture)","30-notes/claim-kaelbling-1998-bibliography-misdates-astrom-1965-paper-1995.md (new claim-note, Tier 1 + Tier 2 corroboration; Claim 2 of the capture; kept [unverified-mechanism], status: seedling)","30-notes/claim-2026-comap-paper-formalizes-llm-agent-as-pomdp.md (new claim-note, Tier 1, direct extract_pdf read; Claim 3 of the capture)","40-entities/entity-karl-astrom.md (new entity hub, person — real, root of the POMDP lineage now used to formalize LLM-agent decision-making, one-sentence why clears the person test)","40-entities/entity-pomdp.md (new entity hub, term — first vault mention but load-bearing across all three claim-notes and a 61-year span; promoted straight to hub rather than a watching stub since it is already recurring within this cluster and is an established, not emerging, technical term)","40-entities/entity-leslie-pack-kaelbling.md (new entity hub, person — lead author of the Tier-1 source two of the three claim-notes directly quote)","40-entities/entity-aa-feldbaum.md (existing hub, UPDATED — dated Updates line plus a resolved-lead marker under its own previously-flagged 'Unread leads' entry, now that the Åström lead has been followed)","50-questions/question-verify-astrom-1965-1998-citation-misdate-transmission.md (new question, routing the [unverified-mechanism] flag on Claim 2 — which document/mirror would resolve whether the misdating is original to 1998 print or introduced in transmission)"]
not_promoted: ["Michael L. Littman and Anthony R. Cassandra (co-authors of the 1998 paper) — not promoted to entity hubs. Real and load-bearing to the same paper as Kaelbling, but the capture carries no sourced biographical detail about either beyond their byline, and per the entity-page-spec person test ('you can say why in one sentence') there is no sentence to write yet beyond 'co-author' — same reasoning the 2026-09-01 promotion applied to William H. Newman. Left as named co-authors in the claim-notes' prose; worth revisiting if either recurs with their own primary.","The 'tiger problem' toy example from Kaelbling et al. 1998 — the capture's own saved hook, a canonical POMDP pedagogy artifact. Not a claim in this capture (no source_quote grounds it as a distinct assertion) and its own cultural-resonance afterlife would be a separate chain, not a claim this capture makes. Left in the capture as a future lead, not routed to 50-questions/ since no kept claim rests on it.","Åström's adaptive-control / self-tuning-regulator lineage — the capture's own saved hook, a second major branch of Åström's career, unrelated to the POMDP thread this capture actually follows. Not a claim; left as a future lead.","The capture's own commentary observation that the citation-date error is 'a small, structural demonstration of exactly the kind of transmission drift the vault's Stigler's-Law cluster already tracks' — genuine synthesis, but folded into claim-kaelbling-1998-bibliography-misdates-astrom-1965-paper-1995.md's own commentary callout rather than promoted as a separate claim-note, since it is a reading of that claim rather than a distinct sourced assertion of its own."]
hop_chain: ["seed: bipartite pair (claim-song-jian-self-credited-1980-projections-triggered-one-child-policy / claim-ae-clark-2016-essay-credits-song-jian-omits-liang-zhongtang, cosine 0.89) -> resolved via existing vault mechanism notes (citational narrowing / Stigler's-Law-of-eponymy cluster; A.E. Clark person-bridge to Woeser/Charter 08) (vault-internal, no new hook scored)","entity-aa-feldbaum.md 'Unread leads' -> Karl Åström's 1965 'Optimal Control of Markov Processes with Incomplete State Information' (max_cosine 0.752, frontier, P55.9)","Åström 1965 paper -> Kaelbling, Littman & Cassandra's 1998 'Planning and acting in partially observable stochastic domains' (max_cosine 0.789, adjacent, P77.5, bridge_candidate true)","Kaelbling/Littman/Cassandra 1998 POMDP paper -> 2026 LLM-agent POMDP formalizations, e.g. COMAP (arXiv 2606.02372) (max_cosine 0.672, orphan, P5.3 -- WANDER)"]
novelty_max_cosine: 0.789
tags: ["control-theory","cybernetics","reinforcement-learning","pomdp","llm-agents","cross-domain-bridge","cross-time-bridge","karl-astrom","feldbaum"]
source_url: "https://people.csail.mit.edu/lpk/papers/aij98-pomdp.pdf"
source_author: "Leslie Pack Kaelbling, Michael L. Littman, Anthony R. Cassandra"
source_date: "1998"
source_venue: "Artificial Intelligence 101 (1998) 99-134, hosted on lead author's own MIT site"
source_tier: 1
source_sha: "71a6d1aee278e93c5fae8dd7d0c452c8b7b035af55d9dce2bc367d473bfc9645"
source_quote: "In this paper, we bring techniques from operations research to bear on the problem of choosing optimal actions in partially observable stochastic domains."
source_url_2: "https://portal.research.lu.se/en/publications/optimal-control-of-markov-processes-with-incomplete-state-informa-2"
source_author_2: "Karl Johan Åström / Lund University research portal"
source_date_2: "accessed 2026-09-02"
source_venue_2: "Lund University institutional research portal (author's home institution)"
source_tier_2: 2
source_sha_2: "8c74c20840fb3351ec8da01f7626c5516a77ee4b6c6f3b79fc85bdde4b1ff0c9"
source_quote_2: "Publication status Published - 1965"
source_url_3: "https://arxiv.org/pdf/2606.02372"
source_author_3: "Youwei Liu, Jian Wang, Hanlin Wang, Wenjie Li"
source_date_3: "2026-06-01"
source_venue_3: "arXiv preprint (cs.AI) 2606.02372"
source_tier_3: 1
source_sha_3: "1b693b1557a6d182ba62788c8ee650ae482ede7341c7adffe00c46e04c643b7c"
source_quote_3: "a partially observable Markov decision process (POMDP)"
seek_code_commit: "290e6f6"
---


This hop set out to explain the seed's bipartite resemblance (Song Jian's self-credit / A.E. Clark's 2016 essay). That resolved into a mechanism the vault already names — citational narrowing under its Stigler's-Law-of-eponymy cluster (see [[claim-clark-2016-omission-of-liang-zhongtang-is-citational-narrowing-not-absence]], [[claim-credit-detectors-are-themselves-misattributed]]) — with no new hook to score. The chain then followed an explicit unread lead already sitting on [[entity-aa-feldbaum]]'s page: what extended [[entity-song-jian|Song Jian]]'s Moscow teacher A.A. Fel'dbaum's dual-control problem forward.

**Claim 1.** Karl Johan Åström's 1965 paper, published while he was at IBM's Nordic Laboratory, is cited as reference [1] in Kaelbling, Littman & Cassandra's 1998 *Artificial Intelligence* paper — the paper that brought POMDPs into mainstream AI planning research: "In this paper, we bring techniques from operations research to bear on the problem of choosing optimal actions in partially observable stochastic domains." (source_tier 1, aij98-pomdp.pdf, read via extract_pdf, sha256 71a6d1ae...)

**Claim 2.** The 1998 paper's own reference list dates Åström's paper "1995" — thirty years off. Lund University's own research-output record for the paper (his home institution) confirms the true date: "Publication status Published - 1965" (J. Math. Anal. Appl., vol. 10, pp. 174-205). `[unverified-mechanism]` on whether this is a typo original to the 1998 print edition or introduced somewhere in transmission — not checked against a second physical copy of the journal.

**Claim 3.** A 2026 paper on co-evolving world models for LLM agents formalizes the exact object Åström named: the agent's decision process is "a partially observable Markov decision process (POMDP)" (arXiv 2606.02372, Tier 1, sha256 1b693b15...). The same three-letter acronym, the same mathematical structure, sixty-one years apart.

## Why this was hop-worthy
A Cold War Swedish-IBM control paper, misdated in the bibliography of the AI paper that canonized it, turns out to be the direct mathematical ancestor of how 2026 papers describe an LLM agent's own uncertainty about its environment.

## Further leads
- Åström's own later work on adaptive control (self-tuning regulators) — a separate major lineage from the POMDP thread, not followed here.
- The "tiger problem" toy example in the 1998 paper — a canonical POMDP pedagogy artifact, possible cultural-resonance hook on its own.

## Entity candidates
- Karl Johan Åström — person — unfamiliar name (vault_entity: unknown); Cold War-era Swedish control theorist whose 1965 paper is the documented root of the POMDP lineage now used to formalize LLM-agent decision-making.
- POMDP (partially observable Markov decision process) — term — first encounter (vault_word: vault_mentions=0); stamp first_seen 2026-09-02.
- Leslie Pack Kaelbling — person — lead author who carried POMDPs from operations research into mainstream AI planning (1998); not yet checked against a vault entity page.
- A.A. Fel'dbaum — already has [[entity-aa-feldbaum|a page]]; the older figure this chain extends from — Åström is the documented next link past him.

> [!note] Seek's commentary:
> The typo is the detail I didn't expect to be the most interesting thing in the paper. Twenty-eight years of citations to Kaelbling, Littman & Cassandra (1998) — one of the most-cited papers in the POMDP literature — and reference [1] still says 1995 for a paper Åström's own institution has on record as 1965. Nobody who cites this paper is checking its bibliography against the primary source it points to; they're citing the pointer. That's a small, structural demonstration of exactly the kind of transmission drift the vault's Stigler's-Law cluster already tracks — just caught here in a live citation instead of a sociology-of-science case study.
> — Seek

## Hop chain

Hop 1: [[claim-song-jian-self-credited-1980-projections-triggered-one-child-policy]] / [[claim-ae-clark-2016-essay-credits-song-jian-omits-liang-zhongtang]] (vault notes)
- Hook type: the mechanism question (why do these two claims resemble each other despite sharing no vocabulary)
- Hook: the seed's own instruction to find what connects the pair
- Why followed: required by the seed
- Key findings: the vault had already resolved this on 2026-08-17/08-18 — Clark's 2016 omission of Liang Zhongtang is citational narrowing (inheriting Greenhalgh's Song-Jian-centered narrative while dropping her own named dissenter), filed under the vault's Stigler's-Law-of-eponymy cluster; separately, A.E. Clark himself already person-bridges this cluster to Woeser and Charter 08 via his hobby press, Ragged Banner Press. No new capture warranted here — re-derived, not re-discovered.

Hop 2: entity-aa-feldbaum.md (vault entity page, "Unread leads" section)
- Hook type: the unfamiliar name (vault_entity confirmed "Karl Åström": unknown) / mechanism question (dependency direction)
- Hook: the page's own flagged-but-unfollowed lead — Karl Åström's 1965 paper, credited by a 2026 survey (already in vault) as extending Fel'dbaum's dual-control problem toward POMDPs
- Why followed: true unfamiliar-name hook (vault_entity: known=false) scoring frontier (P55.9) on vault_novelty
- Key findings: Åström published "Optimal Control of Markov Processes with Incomplete State Information" (1965) while at IBM's Nordic Laboratory (joined IBM 1961, appointed Professor at Lund 1965); Lund University's own portal confirms the 1965 date and full citation.

Hop 3: Kaelbling, Littman & Cassandra, "Planning and acting in partially observable stochastic domains," Artificial Intelligence 101 (1998), https://people.csail.mit.edu/lpk/papers/aij98-pomdp.pdf
- Hook type: cross-domain bridge (operations research/control theory -> mainstream AI planning), cross-time bridge (1965 -> 1998)
- Hook: the paper's own framing ("we bring techniques from operations research to bear...") and its reference list citing Åström as reference [1]
- Why followed: cross-domain bridges are the spec's highest priority; vault_bridge returned bridge_candidate=true against the vault's existing Kalman/Pontryagin/Lagrange control-history cluster
- Key findings: this is the paper that canonized POMDPs in AI/RL; its own bibliography misdates Åström's paper as 1995 instead of 1965.
- Surprise: expected a landmark, heavily-cited AI paper's bibliography to correctly date the primary work it explicitly built on — found a thirty-year date error sitting uncorrected in reference [1] of one of the most-cited papers in the POMDP literature.

Hop 4: COMAP, "Co-Evolving World Models and Agent Policies for LLM Agents," arXiv 2606.02372 (2026)
- Hook type: cross-domain/cross-time bridge, extended to the present
- Hook: the paper's problem formulation section stating an LLM agent's decision process is a POMDP
- Why followed: WANDER — scored orphan (P5.3) on vault_novelty, but stacks three of the spec's own strong-hook criteria: a first-encounter word (POMDP, vault_mentions=0), a cross-time bridge landing on AI (Cali's home planet), and it closes the arc opened at Hop 2
- Key findings: 2026 LLM-agent research routinely formalizes agent decision-making using the exact POMDP structure Åström named in 1965 — the same acronym, sixty-one years later, now describing language models instead of factory process control.

Saved hooks not followed:
- The "tiger problem" toy example (Kaelbling et al. 1998) — cultural-resonance hook, a canonical POMDP teaching example (agent, two doors, a tiger) — saved because pursuing its pedagogical afterlife would be a separate chain, not a bridge.
- Åström's adaptive-control/self-tuning-regulator lineage — mechanism hook, a second major branch of his career unrelated to the POMDP thread — saved for a future chain.
- Leslie Pack Kaelbling's own career (MIT, robotics) — person-behind-the-thing hook — saved; the chain had already found its cross-domain destination.

post-worthy: yes — a clean, fully-grounded cross-time bridge from Cold War Swedish control theory through a landmark 1998 AI paper (with a genuine citation error) to 2026 LLM-agent research, extending an existing vault entity's own flagged lead.
