Capture: Re-sourcing the 2026 inference-vs-training compute share to primary sources (Epoch AI, Stanford HAI AI Index) — replacing Telnyx in V-001/V-002/V-011
Answering the queued remediation task from ruling 7: re-source the figures underlying claim-inference-cost-collapsed-280x (audit V-001) and claim-inference-dominant-ai-compute-2026 (audit V-002/V-011) away from their shared Telnyx citation, direct to Epoch AI and Stanford HAI's AI Index — the two primary venues the remediation ruling named.
Short answer: the cost-collapse claim (V-001) is now cleanly re-sourced — both its headline figures resolve directly to Tier 1 primary pages, word for word. The compute-share claim (V-002/V-011) is not resolved by this research. Neither Epoch AI's nor Stanford HAI's own venues publish the "two-thirds of AI compute is inference in 2026" figure or an equivalent industry-wide measurement. What this session did find, on the same two venues plus one additional Tier 1 arXiv source, is a set of competing primary-adjacent estimates of the training/inference split that disagree with each other by a wide margin and with the "inference dominates" narrative specifically — which is itself informative and is recorded below rather than discarded.
Claim 1 (re-sources V-001): The cost to query a GPT-3.5-equivalent model fell from $20 to $0.07 per million tokens between November 2022 and October 2024 — a 280-fold reduction
Claim type: quantitative Sourcing floor: Tier 1–2 required; achieved Tier 1, replacing the Telnyx relay
"The cost of querying an AI model that scores the equivalent of GPT-3.5 (64.8% accuracy) on MMLU dropped from $20 per million tokens in November 2022 to just $0.07 per million tokens by October 2024 (Gemini-1.5-Flash-8B)—a more than 280-fold reduction in approximately 18 months."
Stanford HAI, "AI Index 2025: State of AI in 10 Charts," 2025-04-07.
Provenance: https://hai.stanford.edu/news/ai-index-2025-state-of-ai-in-10-charts — Tier 1 (Stanford HAI's own research index, own venue, primary data compilation). This is the figure's actual origin — Telnyx's page had relayed it correctly in substance, but citing Telnyx put a Tier 3 layer between the vault and the number for no reason. One correction to the existing claim-note worth flagging at promotion: the specific model hitting $0.07/Mtok is named here as Gemini-1.5-Flash-8B, a detail the existing note omits.
Claim 2 (re-sources V-001): Depending on performance tier/benchmark, annual LLM inference price-decline rates range from roughly 9x to 900x, with a post-2024 median around 200x/year
Claim type: quantitative Sourcing floor: Tier 1–2 required; achieved Tier 1, replacing the Telnyx relay
"ranging from 9x to 900x per year"
"the price to achieve GPT-4's performance on a set of PhD-level science questions fell by 40x per year"
"the median rate increased from 50x per year to 200x per year" (when restricting to data after January 2024)
Ben Cottier, Ben Snodin, David Owen, and Tom Adamczewski, Epoch AI, "LLM inference prices have fallen rapidly but unequally across tasks," 2025-03-12.
Provenance: https://epoch.ai/data-insights/llm-inference-price-trends — Tier 1 (Epoch AI's own data-insight page, named authors, own methodology: log-linear regression on the cheapest model meeting each of six benchmark performance thresholds over three years). This is the actual Epoch AI page the existing claim-note's "9x-900x" figure traces back to; the vault previously had it only via Telnyx's secondhand citation of "Epoch AI." Fetched and confirmed live this session.
Claim 3 (attempted re-source of V-002/V-011; not achieved): Neither Epoch AI nor Stanford HAI's AI Index publishes an industry-wide "two-thirds of AI compute is inference in 2026" figure
Claim type: quantitative (a null result on a quantitative claim) Sourcing floor: Tier 1–2 required for the underlying number; not found on either named primary venue
Both venues named in the remediation task were checked directly for this specific figure:
- Stanford HAI's 2026 AI Index Report (https://hai.stanford.edu/ai-index/2026-ai-index-report) and its Chapter 1: Research and Development sub-page (https://hai.stanford.edu/ai-index/2026-ai-index-report/research-and-development) were fetched directly. Both cover data-center counts, global compute capacity, and energy/environmental footprint, but neither states a training-vs-inference compute split for 2025 or 2026.
- Epoch AI's most topically adjacent recent piece, Josh You's "Frontier labs don't use most AI compute (yet)" (2026-05-21, https://epochai.substack.com/p/frontier-labs-dont-use-most-ai-compute), was fetched and read in full. It discusses the distribution of total AI compute across companies (frontier labs vs. everyone else) but does not state an inference-vs-training percentage split.
Conclusion for this claim: this reinforces, from the opposite direction, what the prior capture (myth-inference-two-thirds-of-compute, via its citation-chain trace to a SambaNova CEO's WEF op-ed) already found — the "two-thirds inference in 2026" figure has no home on either primary venue this remediation task pointed at. It is not merely under-cited; a direct search of the correct primary venues for it comes up empty.
Claim 4 (new primary material, not previously in the vault): One Tier 1 modeling estimate puts training at ~40% of compute in 2024–2026, with an explicit statement that public estimates on this split vary widely
Claim type: quantitative Sourcing floor: Tier 1–2 required; achieved Tier 1
"in 2024 approximately 40% of compute was used for training (including both pre- and post-training), with the remaining share going to model inference and other uses."
"in 2025 and 2026 40% of compute will be used for model training, 30% at the start of 2027 and 20% by the end of 2027."
"There is considerable uncertainty in the estimated present day compute allocations between training, inference and other workloads, with public estimates varying widely."
Iyngkarran Kumar (University of Edinburgh; Centre for the Governance of AI winter fellow) and Sam Manning (Centre for the Governance of AI), "Trends in Frontier AI Model Count: A Forecast to 2028," arXiv:2504.16138.
Read at face value, a 40%-training / ~60%-inference-and-other split for 2025–2026 is directionally consistent with "inference is the majority of compute" — but nowhere near as extreme as "two-thirds inference, up from a third in 2023," and the authors flag their own aggressive 2027 endpoint (20% training) as a scenario they then walked back to a "more balanced 30-70 split" in their baseline forecast. This is a modeled estimate with named authors and disclosed assumptions, not a measurement of actual industry-wide spending — recorded as such.
Provenance: https://arxiv.org/abs/2504.16138 (abstract) and https://arxiv.org/html/2504.16138v1 (full text, fetched directly) — Tier 1.
Claim 5 (new primary material, in direct tension with the vault's existing "inference dominates" framing): A competing modeled estimate — Epoch AI's GATE model, as cited by Kumar & Manning — puts training at 90% of compute in 2024–2025, declining only to 70% by 2026–2028
Claim type: quantitative Sourcing floor: Tier 1–2 required; achieved Tier 1 via the citing paper; the underlying Epoch AI venue itself could not be made to state the figure directly (see caveat)
"Another source for the compute allocations is the Epoch GATE model – see the training-inference split graph."
Kumar & Manning, arXiv:2504.16138, Appendix H ("Results for alternate training compute allocations"), Table 15: the GATE model's training-allocation assumption is given as 90% for 2024–2025, declining to 70% for 2026–2028 — implying only 10–30% of compute goes to inference and other uses, the opposite direction from "inference is two-thirds."
Caveat, recorded rather than smoothed over: this session fetched Epoch AI's own GATE announcement post (https://epoch.ai/blog/announcing-gate, 2025-03-21) and the GATE model's own arXiv paper (https://arxiv.org/abs/2503.04941 and its full HTML text) directly, specifically to verify the 90%/70% figures against Epoch's own words rather than relying on a secondary citation of them. Neither Epoch document states these specific percentages in its prose — the GATE paper explains the training/inference tradeoff framework conceptually but the year-by-year allocation numbers appear to live only in the interactive GATE Playground tool's default scenario (https://epoch.ai/gate), which this session did not directly query. The 90%/70% figures are therefore recorded here on Kumar & Manning's (Tier 1, arXiv, named authors) citation of GATE's output, not on an independently-fetched Epoch AI primary statement of the same number. This clears the Tier 1–2 floor (Kumar & Manning is itself a qualifying Tier 1 source reporting the figure), but promotion should note that the ultimate origin — Epoch's own tool — was checked and did not yield the number directly in this session.
Provenance: https://arxiv.org/html/2504.16138v1 (Kumar & Manning, Appendix H) — Tier 1. Underlying tool not independently confirmed: https://epoch.ai/gate (not queried this session).
Claim 6 (context, same tier logic as Claim 4): OpenAI-specific compute allocation was independently estimated by a named researcher at roughly 40% training in 2024, ~30% external inference deployment, with training's share projected to fall further by 2027
Claim type: quantitative Sourcing floor: Tier 1–2 required; achieved Tier 1
"we estimate that OpenAI used an average of 40% of their compute on training" (2024)
"roughly 30% of their compute on external deployment" (i.e. inference serving external users, 2024)
"All estimates are based on publicly available information which is scarce and uncertain"
Romeo Dean, "Compute Forecast," AI Futures Project (ai-2027.com research supplement), 2025-04.
This is a single-company estimate (OpenAI specifically), not an industry-wide figure, and its own author flags it as built on scarce public information. It sits closer to Claim 4's ~40%-training range than to Claim 5's 90%-training range, for whatever that convergence between two independent named forecasters is worth — but it does not resolve the industry-wide question either.
Provenance: https://ai-2027.com/research/compute-forecast — Tier 1 (named author, own research venue, transparent stated assumptions and explicit uncertainty caveat).
Answering the core question
The remediation task asked for the 2026 inference-vs-training compute share to be re-sourced to Epoch AI and Stanford HAI's AI Index, in place of Telnyx. Splitting the two claim-notes this affects:
V-001 (claim-inference-cost-collapsed-280x) — resolved. Both headline per-token cost figures (the 280x drop, and the 9x-900x annual-decline range) now trace directly to Tier 1 primary pages on the two named venues, fetched and quoted verbatim in this session. The Telnyx citation can be dropped entirely for these two figures; nothing in the re-sourcing changed the substance of either number.
V-002/V-011 (claim-inference-dominant-ai-compute-2026) — not resolved; central question remains [unverified — could not confirm or deny after search]. Direct, targeted searches of both named primary venues (Stanford HAI's 2026 AI Index Report and its R&D chapter; Epoch AI's most relevant recent compute-distribution post) turned up no statement of an industry-wide inference-vs-training compute percentage split for 2026, matching or otherwise. This is the second research pass to reach that null result independently (the first being the prior capture's citation-chain trace of the Deloitte/WEF figure) — two different search strategies against two different sets of venues both failing to locate a primary measurement is itself evidence the figure doesn't have one, rather than a search-quality problem.
What this session adds beyond the null result: two genuinely new Tier 1 sources (Kumar & Manning's arXiv paper, and its citation of Epoch's GATE model) that attempt the same estimate the vendor figure claims to describe, and land in sharply different places from each other — 40%-training-consensus-ish (Claim 4) versus 90%-training (Claim 5) versus a single-company 40%-training data point (Claim 6). None of the three is a measurement; all three are modeled forecasts by named authors with disclosed (and in two cases explicitly flagged as uncertain) assumptions. The spread between them — training compute estimated anywhere from 20% to 90% of the total depending on whose model you use — is more informative than any single number: it shows that "share of compute spent on inference vs. training, industry-wide, in 2026" is not a settled empirical question even among people doing careful primary-source modeling, let alone something a vendor blog post can respond with a single confident percentage.
Recommended framing for promotion: keep claim-inference-cost-collapsed-280x's cost figures, re-sourced to Stanford HAI and Epoch AI directly per Claims 1–2 above (this should let V-001's flag close). Do not close V-002/V-011 on claim-inference-dominant-ai-compute-2026 — the flag should stay, but the note can be strengthened by adding Claims 4–6 as "competing modeled estimates, none of which is a primary measurement and none of which agrees with the others or with the two-thirds figure," which gives readers a clearer picture of the actual state of evidence than either silently keeping the vendor number or silently dropping it. myth-inference-two-thirds-of-compute should be updated to note that this second, independent search (targeted at the two venues the remediation ruling specifically named) also came up empty, and that two competing Tier 1 forecasts diverge from the myth's own figure in opposite directions.