A 2026 paper on co-evolving world models for LLM agents formalizes the agent's decision process as a POMDP — the same structure Åström named in 1965
COMAP, "Co-Evolving World Models and Agent Policies for LLM Agents" (arXiv 2606.02372, 2026), formalizes an LLM agent's own decision process as "a partially observable Markov decision process (POMDP)." This is the identical mathematical object Karl Åström named in his 1965 control-theory paper — see claim-kaelbling-1998-cites-astrom-1965-as-pomdp-origin — and the same acronym Kaelbling, Littman & Cassandra carried into mainstream AI planning research in 1998. Sixty-one years separate Åström's original formulation, developed for industrial process control at IBM's Nordic Laboratory, from a 2026 paper using the same structure to describe a language model's uncertainty about the state of the environment it is acting in.
The 2026 paper's own use of the term is unremarkable within reinforcement-learning-adjacent AI research — POMDPs have been a standard formalism for agent decision-making under partial observability since the 1998 canonization — which is itself the point: the framework has become common vocabulary rather than a cited borrowing, the terminal state of the same kind of transmission this vault tracks elsewhere (claim-karpathy-llm-wiki-gist-canonized-compounding-pattern is a much faster instance of the same "pattern becomes vocabulary" arc, compressed into months rather than decades).
Source
“a partially observable Markov decision process (POMDP)”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-09-02-hop-astrom-pomdp-llm-agents.md, 2026-09-02 (headless) · raw markdown