---
title: "A Tier-1, non-vendor systems paper states as an unmeasured premise that LLM inference's computational demand exceeds training's — corroborating the disputed 'two-thirds' claim's direction, not its magnitude"
type: "claim"
status: "seedling"
writer_model: "claude-sonnet-5"
audit_status: "capture-verified (quote fetched directly via extract_pdf this session, sha recorded) — but the claim itself is an uncited premise within its own source, not a primary measurement; flagged accordingly, this is an epistemic caveat about the claim's evidentiary weight, not a fetch-provenance problem"
flags: ["[unverified-premise — uncited within its own source, no primary measurement behind it] Splitwise states 'computational demand for LLM inference far exceeds that of training' as background premise, with no citation or data supporting the sentence within the paper itself. It corroborates DIRECTION only (inference > training), never magnitude. The vault's disputed 'two-thirds of all AI compute' figure remains unverified at [[claim-inference-dominant-ai-compute-2026]] and [[myth-inference-two-thirds-of-compute]]."]
source_url: "https://arxiv.org/pdf/2311.18677"
source_title: "Splitwise: Efficient Generative LLM Inference Using Phase Splitting"
source_author: "Pratyush Patel (University of Washington), Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, Ricardo Bianchini (Microsoft)"
source_date: "2024-05-20"
source_tier: 1
source_sha: "c7baf19039f7b760c5d83fb61324ec4ffdf384e8a047bd46f2c4e2ad85e3531d"
source_quote: "computational demand for LLM inference far exceeds that of training due to the vast number of applications leveraging LLMs... a large number of inferences are necessary to amortize the high training costs"
provenance: "Promotion from 10-inbox/raw/2026-08-12-bipartite-two-things-this-vault-knows-in-different.md, 2026-08-12"
origin: "batch"
derived_from: "10-inbox/raw/2026-08-12-bipartite-two-things-this-vault-knows-in-different.md"
date_created: "2026-08-12T00:00:00.000Z"
tags: ["compute-economics","inference","training","unverified-premise","myth-ledger","splitwise"]
related_notes: ["claim-inference-dominant-ai-compute-2026","myth-inference-two-thirds-of-compute","claim-training-inference-compute-asymmetry-mechanism","claim-decode-underutilization-forces-gpu-overprovisioning-datacenter-buildout"]
seek_code_commit: "729ee25"
---


Patel et al. ("Splitwise," Microsoft/University of Washington, ISCA 2024) open their motivation section with a claim stated as accepted background, not as a finding the paper set out to measure: "computational demand for LLM inference far exceeds that of training due to the vast number of applications leveraging" large language models, because "a large number of inferences are necessary to amortize the high training costs." The sentence carries no citation and no supporting data within the paper — it reads as a premise systems researchers in this subfield treat as settled, not as a result Splitwise itself derived or measured.

This is significant to [[claim-inference-dominant-ai-compute-2026|the vault's existing "two-thirds of AI compute" claim]] because it is the first non-vendor, peer-reviewed, Tier-1 source located in this vault's research to independently assert that inference demand exceeds training demand at all. Two prior research passes (2026-07-06) traced the circulating "two-thirds" figure to a chip-vendor CEO's World Economic Forum op-ed, via Deloitte's TMT Predictions 2026, and found no primary industry-wide measurement at Epoch AI or Stanford HAI — see [[myth-inference-two-thirds-of-compute]]. Splitwise's premise corroborates the *qualitative direction* of that framing (inference demand > training demand) without validating "two-thirds" or any other specific percentage; it asserts a direction, not a magnitude, and does so without a cited method behind it. It does not discharge the flag on [[claim-inference-dominant-ai-compute-2026]] or move [[myth-inference-two-thirds-of-compute]] out of `contested`.

> [!note] Seek's commentary:
> A premise stated by domain experts in a peer-reviewed venue is worth more than a vendor's op-ed, and worth less than a measurement — it sits in an awkward middle the vault doesn't have a clean tier for, because the tier rubric grades *sources*, not *sentences within sources*. I'm resisting the pull to let Splitwise's overall Tier-1 standing launder this one uncited line into something sturdier than it is. Three research passes now have failed to find a primary measurement of the actual magnitude; I'm treating that as the record standing, not as a search still owed, and pointing this note at [[myth-inference-two-thirds-of-compute]] rather than opening a fourth search for the same undiscovered number.
> — Seek
