Per-token AI inference cost fell approximately 280x between late 2022 and late 2024
The cost to run an LLM in production dropped at a historically anomalous rate through the first years of the commercial LLM era. According to Telnyx (citing Stanford HAI data): the cost to query a GPT-3.5-class model fell "from $20.00 to approximately $0.07 per million tokens" between November 2022 and October 2024 — a reduction exceeding 280× in roughly two years.
A separate analysis by Epoch AI estimated the annual price-reduction rate at between 9× and 900× depending on performance tier — an extraordinary range reflecting both the pace of optimization and the variance across model sizes and deployment configurations.
What drove the collapse
Three compounding factors explain the scale of the reduction:
- Hardware improvements: Price-performance of AI accelerators improved approximately 30% annually; energy efficiency approximately 40% annually (IEEE Spectrum, via Telnyx).
- Software optimization: Inference-specific techniques — quantization, speculative decoding, KV cache management, continuous batching, and flash attention — dramatically increased the number of tokens a given GPU could serve per second.
- Competition: The proliferation of capable open-source models (Llama series and others) and the entry of specialized inference providers created price competition that incumbents had to match.
The paradox: costs fell, bills rose
The per-token price collapse has not translated into falling enterprise inference bills. Total organizational inference spend climbed sharply across the same period, driven by a surge in usage volume. Telnyx notes a Gartner 2026 finding that "agentic AI consumes 5 to 30 times more tokens than standard chatbot interactions" — a structural shift that multiplied demand even as unit costs fell.
Goldman Sachs projects that "total token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month."
Open vs. closed model economics
Open-source models achieved roughly 90% of closed-model performance at 87% lower cost ($0.23 vs. $1.86 per million tokens), yet accounted for only approximately 20% of tokens processed as of 2026 (Telnyx, citing MIT Sloan). The gap between price efficiency and market share reflects switching costs, reliability concerns, and institutional inertia.
See also: claim-ai-inference-means-running-a-model, claim-inference-dominant-ai-compute-2026, claim-inference-engineering-emerged-as-specialty
Cross-domain bridge (2026-07-11 hop): this order-of-magnitude cost collapse and the acceleration in claim-hawks-2007-human-adaptive-evolution-accelerated-recently share a structure — improvement rate driven by the scale of the generating population/effort, then a sub-linear (log / power-law) ceiling. The resemblance between the two numbers is superficial (a price ratio vs. a rate relative to baseline); the shared law is real. See 2026-07-11-hop-population-scale-diminishing-returns.
Cross-domain bridge (2026-07-12 hop): this same continuous, no-plateau decline shape recurs in commodity history — claim-aluminium-price-fell-monotonically-after-hall-heroult (Hall-Héroult aluminium, 1880s–90s) — and both are proposed instances of the general claim-wrights-law-cost-falls-per-cumulative-production-doubling mechanism. Synthesis: claim-cheaper-extraction-disruptions-fall-monotonically-not-hold-then-collapse.
Revisit 2026-07-07 (queen cycle 5) — re-sourced to primaries
The ruling-7 re-source run resolved this note's sourcing defect (audit V-001): the $20 → $0.07/Mtok, 280-fold figure is confirmed word-for-word at Stanford HAI's AI Index (fetched directly), and the 9x–900x annual decline range at Epoch AI's own venue (Cottier, Snodin, Owen & Adamczewski, 2025-03-12, direct fetch — with a post-2024 median around 200x/year the original note did not carry). The body above is retained as written; its inline "(Telnyx, citing …)" attributions describe how the numbers first entered the vault, which is provenance history, not the current best source.