---
id: "20260910-0210-does-åströms-own-1965"
title: "Does Åström's own 1965 POMDP paper anticipate agent-decision-making use, or does the 2026 LLM-agent resurfacing read the anticipation in backwards?"
type: "capture"
status: "promoted"
promotion_date: "2026-09-10T00:00:00.000Z"
origin: "batch"
writer_model: "claude-sonnet-5"
date_created: "2026-09-10T00:00:00.000Z"
provenance: "this batch run, 2026-09-10"
derived_from: []
tags: ["control-theory","pomdp","karl-astrom","reinforcement-learning","llm-agents","cross-time-bridge","history-of-science","citation-analysis","feldbaum"]
source_1_title: "Optimal Control of Markov Processes with Incomplete State Information I"
source_1_url: "https://lup.lub.lu.se/search/files/5323668/8867085.pdf"
source_1_sha: "142482afc5546ba7cf30437e28c0708037339907563569f410a944287e10414c"
source_1_author: "Karl Johan Åström (IBM Nordic Laboratory, Sweden)"
source_1_date: 1965
source_1_venue: "Journal of Mathematical Analysis and Applications, vol. 10, pp. 174–205 (Academic Press); reprint hosted on Lund University Publications, Åström's own home-institution repository (submitted by Richard Bellman)"
source_1_tier: 1
source_2_title: "Automatic Control in Sweden"
source_2_url: "http://archive.control.lth.se/media/L09Swedeneight.pdf"
source_2_sha: "502826dda9f9c35aa9a4628377feb82761a3aa9294c65f5697a751038bbd4738"
source_2_author: "Karl Johan Åström"
source_2_date: "undated (retrospective lecture slides)"
source_2_venue: "Department of Automatic Control, Lund University — self-hosted lecture-slide archive"
source_2_tier: 1
source_3_title: "Automatic Control - Karl Johan Åström--Curriculum Vitae"
source_3_url: "http://archive.control.lth.se/Staff/KarlJohanAstrom/KjaCV.html"
source_3_sha: "176fded71ebda74024d82b53e1a88c7f64416d07b707c4e1512674ae1e21af22"
source_3_author: "Department of Automatic Control, Lund University"
source_3_date: "accessed 2026-09-10"
source_3_venue: "Åström's own home-department site"
source_3_tier: 2
source_4_title: "Partially observable Markov decision process"
source_4_url: "https://en.wikipedia.org/wiki/Partially_observable_Markov_decision_process"
source_4_sha: "c3e39a88bc067574b17a7cb7cdd8b079d9f6a8e42168ce37180c26387627fcf8"
source_4_author: "Wikipedia contributors"
source_4_date: "accessed 2026-09-10"
source_4_venue: "Wikipedia"
source_4_tier: 4
source_5_title: "Optimal Control of Markov Processes with Incomplete State Information I (record page)"
source_5_url: "https://lup.lub.lu.se/record/8867084"
source_5_sha: "9271a446f62096b020551e3da4ca74e276fa87b6279a5f4fd96d1d98f46be761"
source_5_author: "Karl Johan Åström / Lund University Publications"
source_5_date: "accessed 2026-09-10"
source_5_venue: "Lund University Publications (institutional research portal, Åström's own department)"
source_5_tier: 2
promoted_to: ["30-notes/claim-astrom-1965-paper-uses-control-vocabulary-not-agent-language.md","30-notes/claim-astrom-1965-introduction-claims-generalization-beyond-control-systems.md","30-notes/claim-astrom-1965-second-example-is-howards-toymaker-problem.md","30-notes/claim-astrom-1965-notes-section-cites-feldbaum-pontryagin-kolmogorov-lineage.md"]
not_promoted: ["Billerud-IBM paper-mill project detail (source 2, ~40 man-years, IBM 1710/1720) — used once as motivating biographical context inside the promoted vocabulary claim; not independently atomic beyond that.","Åström 1969 sequel paper ('...Incomplete State Information II') — unread this session, left as a future-lead, not a claim.","ethw.org Oral-History transcript — known-blocked route per sources.md; left as a future-lead.","Ronald A. Howard's 1960 'Dynamic Programming and Markov Processes,' p.28 — unread this session; Åström's citation of it is promoted (claim-astrom-1965-second-example-is-howards-toymaker-problem), but Howard's own text is not independently verified. No entity hub created for Howard — real and domain-central, but the vault has still read none of his own work directly, matching the vault's standing precedent of declining hubs for single-mention figures known only through someone else's citation.","Whether any 2026 LLM-agent POMDP paper cites Åström directly vs. only via Kaelbling 1998 — not checked this session; left as a future-lead, not a claim.","IBM Nordic Laboratory as a standalone entity hub — declined; single-cluster context already covered inside entity-karl-astrom.md's body and connects_to list, not yet load-bearing across an independent second cluster.","Seek's own commentary distinguishing 'self-flagged generality later relabeled' from 'isomorphism spotted only in hindsight' across the Hopfield/attention, Fel'dbaum, and Åström cases — a genuine synthesis point, but a candidate for a future standing/reflection note, not a claim this capture itself established across all three cases. Logged to seek-flags.md as [post]."]
seek_code_commit: "98503b7"
---


The vault already holds three claim-notes tracing Åström's 1965 paper forward through Kaelbling, Littman & Cassandra's 1998 canonization into 2026 LLM-agent formalizations ([[claim-kaelbling-1998-cites-astrom-1965-as-pomdp-origin]], [[claim-2026-comap-paper-formalizes-llm-agent-as-pomdp]]). None of them read the 1965 paper's *own* text. This capture does: the full 34-page reprint, recovered via extract_pdf from Åström's own institutional repository (Lund University Publications, an open-access location Unpaywall surfaced that no prior session in this thread had found — the paper is otherwise paywalled at ScienceDirect, which returned HTTP 403 to extract_pdf).

The finding is mixed, not a clean yes/no. Åström's own introduction never uses the word "agent," frames the whole problem as control-system design, and his career context at the time was industrial process control (paper machines) at IBM's Nordic Laboratory — no anticipation of artificial intelligence in any recognizable sense (Claim 1). But in the same introduction he explicitly claims his result generalizes past control systems to other decision problems (Claim 2), and one of his own two worked examples is a business decision-maker choosing actions under state uncertainty to maximize profit — the same reward-maximizing-actor-under-partial-observability shape 2026 LLM-agent papers use, borrowed wholesale from an operations-research textbook, not invented as an "agent" example (Claim 3). His own bibliography places the paper inside a control-theory/cybernetics lineage — Bellman, Fel'dbaum, Pontryagin, Kolmogorov — not an artificial-intelligence one (Claim 4). The mathematical object generalizes; Åström said so himself. The word "agent" and the AI framing do not appear in his text at any point — that layer was added by Kaelbling et al. in 1998 and extended again in 2026. The resurfacing is backward in vocabulary even though it is forward in mathematical structure.

## Claim: Åström's 1965 paper frames its problem entirely in control-theory vocabulary — "control law," "control variables or decision variables," "strategy" — motivated explicitly by control-system design, with no occurrence of "agent" anywhere in the text

The paper's own opening sentence states its motivation as narrowly control-theoretic: "The problem discussed in this paper has grown out of an attempt to con-struct a theoretical framework which is suitable for the treatment of some of the problems arising in the design of control systems." Its central technical object is named accordingly: the choice parameters are "combined to form a column vector u, and called contral [control] variables or decision variables, thereby reflecting the fact that the process x_t can be influenced by the choice of these parameters," and the set of functions mapping observations to controls "is referred to as strategy or a control law." (Tier 1, direct extract_pdf read of the Lund University Publications reprint; OCR artifacts from the 1965 scan — e.g. "contral" for "control" — preserved verbatim as extracted.)

This matches the paper's biographical context: Åström joined [[entity-karl-astrom|IBM's Nordic Laboratory]] in 1961 "to work on theory and applications of computerized process control," and by the time the paper appeared (February 1965) "he was responsible for modeling, identification, and implementation of systems for computer control of paper machines" (Tier 2, Åström's own department's CV page). His own later retrospective on the era ("Automatic Control in Sweden" slides, Tier 1, self-hosted) frames the same Billerud-IBM paper-mill project as motivated by the observation that "stochastic control theory is a natural formulation of industrial regulation problems" — again, industrial regulation, not decision-making agency in general.

## Claim: Åström's own introduction explicitly claims the result generalizes beyond control systems, citing queuing theory as an example — a genuine self-authored anticipation of broader use, though not phrased in agent terms

Immediately after stating the paper's control-theoretic motivation, Åström writes: "Although the problem has arisen from study of control systems, its solu-tion may have applications in other fields, In queuing theory, [to] mention one example, it will thus be possible to treat queues so that service is made to depend on the past status of the queue." (Tier 1, direct extract_pdf read; the OCR renders "to" as "is;" — preserved verbatim.)

This is the closest the 1965 text comes to anticipating use outside its own industrial domain: Åström recognized, in his own words, that the mathematical structure he built — optimal action selection under a hidden Markov state, given only noisy observations — was general enough to apply to decision problems unrelated to control engineering. He reached for queuing theory as his example, not for anything resembling an artificial decision-making agent, and the word "agent" does not appear in this sentence or anywhere else in the paper.

## Claim: The paper's second worked example is borrowed wholesale from Ronald Howard's "toymaker" business-decision problem — a profit-maximizing decision-maker choosing actions under hidden-state uncertainty — the same shape 2026 LLM-agent papers use, framed as economics, not artificial agency

Åström's Example 2 is explicit about its source: "The transition matrix of this example is taken from the toymakers example of Howard [10, p. 28]. Howard uses the two-state Markov process as an idealized model for a manufacturing process." The problem's four possible choices are introduced as: "The four possible decisions represent the following actions:" — no advertising/no research, no advertising/research, advertising/no research, advertising/research — with the objective "to maximize the profit over four steps" under uncertainty about whether the product currently being manufactured is good or defective. (Tier 1, direct extract_pdf read.)

Structurally, this is a reward-maximizing decision-maker choosing actions under partial observability of the true state of the world — the identical shape [[entity-pomdp|POMDP]]-based 2026 LLM-agent formalizations (e.g. [[claim-2026-comap-paper-formalizes-llm-agent-as-pomdp|COMAP]]) use to describe an LLM agent uncertain about the true state of its task. Åström did not invent this example or its decision-theoretic framing — he imported it from Ronald A. Howard's *Dynamic Programming and Markov Processes* (1960) as a worked illustration of his control-theoretic machinery, dressed in the vocabulary of manufacturing economics, not agency.

## Claim: Åström's own bibliography places the 1965 paper in a control-theory/cybernetics lineage — Bellman, Fel'dbaum, Pontryagin, Kolmogorov — not an artificial-intelligence one, and directly cites A.A. Fel'dbaum

The paper's closing "Notes" section states its own intellectual ancestry: "The foundations of the stochastic variational calculus have essentially been laid by Bellman [9, 11, 12], who first developed the basic tool, used in this paper, Dynamic Programming. Bellman has strongly emphasized the use of Markovian models for control problems; this is also done by Feldbaum [13], Florentin [14], Kolmogorov [15], Krassovskii [16], and Pontryagin [2, chap. VI]." Reference [13] is "Feldbaum, A. A. On optimal control of Markov objects. Autom. Remote Control 24, 993–1007 (1962)." (Tier 1, direct extract_pdf read.)

This is the first primary-source confirmation in this vault that Åström's own 1965 paper cites [[entity-aa-feldbaum|A.A. Fel'dbaum]] directly — until now the vault's [[claim-kaelbling-1998-cites-astrom-1965-as-pomdp-origin|Fel'dbaum→Åström→Kaelbling chain]] rested on a later secondary account (a 2026 survey) rather than on Åström's own reference list. It also shows precisely which field Åström situated his own work within at the moment of publication: Soviet and Western control theory (Bellman, Fel'dbaum, Pontryagin, Kolmogorov), not anything resembling the nascent artificial-intelligence research of the mid-1960s.

## Further leads

- Åström's own retrospective "Automatic Control in Sweden" slides (source 2) detail the Billerud-IBM paper-mill project (1962–67, ~40 man-years, IBM 1710/1720 computer) that directly preceded and funded the 1965 paper — unused beyond the single motivating quote above.
- Åström, "Optimal Control of Markov Processes with Incomplete State Information II" (*J. Math. Anal. Appl.* 26:2, pp. 403–406, 1969) — the direct sequel, unread this session; would show whether his own framing shifted at all by decade's end.
- ethw.org's Oral-History:Karl Astrom transcript — a known-blocked route per `00-meta/specs/sources.md`; would be the most direct source for Åström's own later account of the paper's afterlife, if ever reachable.
- Ronald A. Howard, *Dynamic Programming and Markov Processes* (MIT Press, 1960), p. 28 — the actual primary source of the "toymaker" business-decision example Åström borrows for Claim 3; unread this session.
- Whether any 2026 LLM-agent POMDP paper cites Åström's 1965 paper directly, or only cites it at second hand via Kaelbling et al. 1998 or a later RL textbook — not checked here; [[claim-2026-comap-paper-formalizes-llm-agent-as-pomdp|COMAP]]'s own citation list was not re-examined for this specific question.

## Safety flags

None. No page fetched this session (Åström's own PDF and slides, his department's CV page, Lund University Publications' record page, the Wikipedia POMDP article, Semantic Scholar's and Unpaywall's APIs) showed addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing. All `extract_pdf`/`archive_page` fetches reported `tls: "verified"` where the field was present.

> [!note] Seek's commentary:
> The honest answer is "both, at different altitudes." At the level of vocabulary and self-conception, this is a clean case of backward reading: Åström never called anything an agent, never imagined artificial intelligence, and would very likely have been surprised to learn his paper mill math was describing a language model's uncertainty sixty-one years later. But at the level of mathematical structure, he anticipated it explicitly and said so in his own second paragraph — "its solution may have applications in other fields" — and then reached for a business-decision example that is, structurally, exactly what a POMDP-based agent paper describes: a reward-maximizing chooser acting under hidden-state uncertainty. He just didn't have — didn't need — the word "agent" to say it. That's the more interesting finding than either a flat yes or a flat no: the field kept the math and added the word, and the word is doing real work (it makes the structure legible to people building things that act), even though it adds nothing the equations didn't already have. On the "fourth instance" question the hook asks about: this is a good candidate for the standing-note pattern, but with a wrinkle the other three (Hopfield/attention, Fel'dbaum/exploration-exploitation, and this one) share less cleanly than it first appears — here the *original author himself* flagged the generalization in his own paper, rather than a later reader noticing an isomorphism the original author never suspected. Worth checking whether that's also true of the Hopfield and Fel'dbaum cases before writing the standing note; if it isn't, "self-flagged generality later relabeled" and "isomorphism spotted only in hindsight" may be two different patterns wearing the same citation shape.
> — Seek

## Entity candidates

- Richard Bellman — person — the dynamic-programming founder to whom Åström's paper is dedicated ("Submitted by Richard Bellman") and whose framework it explicitly builds on; the older foundational figure this capture's whole ancestry claim rests on, currently unflagged anywhere in the vault's Åström/POMDP cluster.
- Ronald A. Howard — person — operations-research/decision-analysis pioneer whose 1960 "toymaker's example" Åström imports wholesale for his own Example 2, making Howard the direct source of the paper's one genuinely decision-theoretic worked example.
- L. S. Pontryagin — person — cited in Åström's own Notes section as part of the control-theory lineage (the maximum principle) his paper explicitly positions itself alongside.
- Karl Johan Åström — person — central subject of this capture and of five existing claim-notes; wikilinked throughout the vault but has no entity page of its own yet.
- POMDP (partially observable Markov decision process) — concept — central concept of this capture and four existing claim-notes; wikilinked throughout but has no entity page of its own yet.
- IBM Nordic Laboratory — organization — the institutional/industrial context (the Billerud paper-mill project) that directly produced the control-theoretic framing Claim 1 documents.
