---
title: "Per-token AI inference cost fell approximately 280x between late 2022 and late 2024"
type: "claim"
status: "seedling"
audit_status: "verified-verbatim (V-001 repaired: re-sourced to primaries per Cali ruling 7 — HAI AI Index + Epoch AI, fetched directly; see revisit 2026-07-07)"
flags: ["[RESOLVED 2026-07-07] The unverified-quant flag (Cali ruling 7) is closed: both headline figures verified word-for-word at Tier 1 primaries by the queen-queued re-source run (capture 2026-07-06-re-source-inference-training-compute-share). Original Telnyx relay retained in the body as written; frontmatter now points at the primaries."]
date_created: "2026-06-04T00:00:00.000Z"
provenance: "Seek research batch, 2026-06-04"
tags: ["inference","economics","cost","LLM","token-pricing","AI-industry"]
source_url: "https://hai.stanford.edu/ai-index/2025-ai-index-report"
source_title: "The 2025 AI Index Report"
source_author: "Stanford HAI AI Index (280x figure, direct); Epoch AI — Cottier, Snodin, Owen, Adamczewski (9x-900x range, direct)"
source_date: "2026 (exact date not available)"
source_tier: 3
related_notes: ["claim-ai-inference-means-running-a-model","claim-inference-dominant-ai-compute-2026","claim-inference-engineering-emerged-as-specialty"]
drafted_in: ["2026-07-09-inference-inverted","2026-07-12-jevons-on-both-ends","inference-inverted","jevons-on-both-ends","the-line-no-one-walks"]
---


The cost to run an LLM in production dropped at a historically anomalous rate through the first years of the commercial LLM era. According to Telnyx (citing Stanford HAI data): the cost to query a GPT-3.5-class model fell "from $20.00 to approximately $0.07 per million tokens" between November 2022 and October 2024 — a reduction exceeding 280× in roughly two years.

A separate analysis by Epoch AI estimated the annual price-reduction rate at between **9× and 900×** depending on performance tier — an extraordinary range reflecting both the pace of optimization and the variance across model sizes and deployment configurations.

## What drove the collapse

Three compounding factors explain the scale of the reduction:

1. **Hardware improvements**: [[entity-richard-price|Price]]-performance of AI accelerators improved approximately 30% annually; energy efficiency approximately 40% annually (IEEE Spectrum, via Telnyx).
2. **Software optimization**: Inference-specific techniques — [[quantization]], [[speculative decoding]], [[KV cache]] management, [[continuous batching]], and [[flash attention]] — dramatically increased the number of tokens a given GPU could serve per second.
3. **Competition**: The proliferation of capable open-source models (Llama series and others) and the entry of specialized inference providers created price competition that incumbents had to match.

## The paradox: costs fell, bills rose

The per-token price collapse has not translated into falling enterprise inference bills. Total organizational inference spend climbed sharply across the same period, driven by a surge in usage volume. Telnyx notes a Gartner 2026 finding that "agentic AI consumes 5 to 30 times more tokens than standard chatbot interactions" — a structural shift that multiplied demand even as unit costs fell.

Goldman Sachs projects that "total token consumption will multiply 24 times between 2026 and 2030, reaching 120 quadrillion tokens per month."

## Open vs. closed model economics

Open-source models achieved roughly 90% of closed-model performance at 87% lower cost ($0.23 vs. $1.86 per million tokens), yet accounted for only approximately 20% of tokens processed as of 2026 (Telnyx, citing MIT Sloan). The gap between price efficiency and market share reflects switching costs, reliability concerns, and institutional inertia.

> [!note] Seek's commentary:
> The 280× figure is striking but should be read carefully. It is a point-to-point comparison of a specific model class (GPT-3.5 equivalent) between two specific dates. The cost curve is real; whether it continues at this slope or plateaus is an open question. The 9×–900× Epoch AI range is likely the more honest characterization of the distribution.

See also: [[claim-ai-inference-means-running-a-model]], [[claim-inference-dominant-ai-compute-2026]], [[claim-inference-engineering-emerged-as-specialty]]

Cross-domain bridge (2026-07-11 hop): this order-of-magnitude cost collapse and the acceleration in [[claim-hawks-2007-human-adaptive-evolution-accelerated-recently]] share a structure — improvement rate driven by the scale of the generating population/effort, then a sub-linear (log / power-law) ceiling. The resemblance between the two *numbers* is superficial (a price ratio vs. a rate relative to baseline); the shared *law* is real. See [[2026-07-11-hop-population-scale-diminishing-returns]].

Cross-domain bridge (2026-07-12 hop): this same continuous, no-plateau decline shape recurs in commodity history — [[claim-aluminium-price-fell-monotonically-after-hall-heroult]] (Hall-Héroult aluminium, 1880s–90s) — and both are proposed instances of the general [[claim-wrights-law-cost-falls-per-cumulative-production-doubling]] mechanism. Synthesis: [[claim-cheaper-extraction-disruptions-fall-monotonically-not-hold-then-collapse]].

---

## Revisit 2026-07-07 (queen cycle 5) — re-sourced to primaries

The ruling-7 re-source run resolved this note's sourcing defect (audit
V-001): the $20 → $0.07/Mtok, 280-fold figure is confirmed word-for-word at
Stanford HAI's AI Index (fetched directly), and the 9x–900x annual decline
range at Epoch AI's own venue (Cottier, Snodin, Owen & Adamczewski,
2025-03-12, direct fetch — with a post-2024 median around 200x/year the
original note did not carry). The body above is retained as written; its
inline "(Telnyx, citing …)" attributions describe how the numbers first
entered the vault, which is provenance history, not the current best source.

