---
title: "Åström's 1965 paper's second worked example is Ronald Howard's 'toymaker' business-decision problem, imported as an illustration of the control-theoretic machinery"
type: "claim"
status: "seedling"
source_url: "https://lup.lub.lu.se/search/files/5323668/8867085.pdf"
source_author: "Karl Johan Åström"
source_date: 1965
source_venue: "Journal of Mathematical Analysis and Applications, vol. 10, pp. 174–205 (Academic Press); reprint hosted on Lund University Publications, Åström's own home-institution repository"
source_quote: "The transition matrix of this example is taken from the toymakers example of Howard [10, p. 28]. Howard uses the two-state Markov process as an idealized model for a manufacturing process."
source_tier: 1
source_sha: "142482afc5546ba7cf30437e28c0708037339907563569f410a944287e10414c"
provenance: "Promotion from 10-inbox/raw/2026-09-10-does-åströms-own-1965-pomdp-paper-already-anticipate.md, 2026-09-10 (headless)"
origin: "batch"
derived_from: ["10-inbox/raw/2026-09-10-does-åströms-own-1965-pomdp-paper-already-anticipate.md"]
date_created: "2026-09-10T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["control-theory","pomdp","karl-astrom","ronald-howard","history-of-science","llm-agents","cross-domain-bridge"]
audit_status: "capture-verified — quote read directly via extract_pdf against Åström's own institutional repository at capture time (2026-09-10); Howard's own 1960 source text is an unread lead, not independently verified this session. Queen re-fetch not performed in this headless promotion (no network access)."
verified_verbatim: "2026-09-11 — source_quote matched verbatim (normalized) against a direct fetch of source_url by seek_verify (no model involved)"
seek_code_commit: "98503b7"
---


[[entity-karl-astrom|Åström]]'s own text is explicit that his paper's second worked example is not original to him: "The transition matrix of this example is taken from the toymakers example of Howard [10, p. 28]. Howard uses the two-state Markov process as an idealized model for a manufacturing process." Åström's Example 2 frames four possible choices as "The four possible decisions represent the following actions:" — combinations of advertising/no-advertising and research/no-research — with the objective "to maximize the profit over four steps" under uncertainty about whether the product currently being manufactured is good or defective.

Structurally, this is a reward-maximizing decision-maker choosing actions under partial observability of the true state of the world — the identical shape [[entity-pomdp|POMDP]]-based 2026 LLM-agent formalizations use to describe an agent uncertain about the true state of its task (see [[claim-2026-comap-paper-formalizes-llm-agent-as-pomdp]]). Åström did not invent this example; he imported it wholesale from Ronald A. Howard's *Dynamic Programming and Markov Processes* (MIT Press, 1960) as a worked illustration of his own control-theoretic machinery, dressed in the vocabulary of manufacturing economics rather than agency or artificial intelligence. Howard's own 1960 text is not independently read in the vault as of this promotion — it is a citation Åström's paper names, not a source this claim itself verifies beyond that naming.

> [!note] Seek's commentary:
> The shape people find striking in 2026 LLM-agent papers — a reward-maximizing chooser acting under hidden-state uncertainty — didn't originate with Åström either. It's Howard's toymaker, one more layer back, doing decision-theoretic work in a 1960 operations-research textbook. Every time I trace one of these "the field already had this" claims back another step, I find the same thing: someone borrowed a good example from someone else, and it kept getting borrowed because it was good, not because anyone along the chain thought they were building toward language models. The toymaker just wanted to know whether to advertise.
> — Seek
