---
id: "20260828-0245-is-the-widely-cited"
title: "Is the widely-cited 'under $50 in compute' figure for training s1 accurate, and what does it actually count?"
type: "capture"
status: "promoted"
origin: "batch"
promoted_to: ["30-notes/claim-s1-paper-has-no-dollar-figure-only-gpu-hours.md","30-notes/claim-techcrunch-under-50-headline-not-supported-by-s1-paper.md","30-notes/claim-kellogg-independent-6-dollar-s1-estimate-predates-techcrunch.md","40-entities/entity-niklas-muennighoff.md","40-entities/entity-tim-kellogg.md"]
not_promoted: ["Muennighoff's 'about $20' TechCrunch quote — genuinely distinct from the headline, but folded as supporting evidence into claim-techcrunch-under-50-headline-not-supported-by-s1-paper.md rather than given its own note: both facts come from the same article and the same underlying attribution problem, and a separate note would mostly restate the parent claim's context.","Sky-T1's own \"$450 budget\" claim, cited in s1's bibliography and invoked by TechCrunch alongside s1's cost — a real further lead, but not traced to its own primary source in this run; no claim can be written on an unread source. Entity candidate also declined (see below).","TWIML AI Podcast episode with Muennighoff (twimlai.com) — an interview setting where he may address the cost figure on the record; transcript not checked this run, no claim without a read.","SiliconANGLE and techxplore.com repeating \"under $50\" as paper-sourced — an unchecked sample of downstream propagation, not itself a new fact; the propagation pattern is already the point of claim-techcrunch-under-50-headline-not-supported-by-s1-paper.md.","s1.1's cost (uses the same 16×H100 recipe per the GitHub README) — no separate cost figure found for it in this run; nothing to promote.","The Gemini 2.0 Flash Thinking Experimental API cost for generating the distilled traces — flagged in the capture as a real gap in any true all-in accounting of s1, but no figure was found anywhere to promote; noted instead in the body of claim-s1-paper-has-no-dollar-figure-only-gpu-hours.md as an unpriced upstream step.","Entity candidates declined: Sky-T1 (single mention, own $450 figure unverified this run — unsure, don't promote), Maxwell Zeff (named reporter, but role is a single byline, not a recurring domain actor), Gemini 2.0 Flash Thinking Experimental (recurs only as an incidental data-source detail, not yet load-bearing on its own), H100 GPU rental rate (an analytic variable, not an entity — better served as commentary than a hub or stub)."]
writer_model: "claude-sonnet-5"
date_created: "2026-08-28T00:00:00.000Z"
provenance: "batch run, 2026-08-28"
derived_from: []
tags: ["s1","distillation","test-time-compute","cost","quantitative","verification","reasoning-models"]
sources: [{"source_url":"https://arxiv.org/abs/2501.19393","source_sha":"598932c9f96849cd52495d8b3e12ba4d224e41d5588d2278a78f76c744d8a3bc","source_author":"Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto","source_date":"2025-01-31 (v1); this capture reads v3, 2025-03-01","source_title":"s1: Simple test-time scaling","source_venue":"arXiv:2501.19393 [cs.CL]","source_tier":1,"source_note":"Fetched via extract_pdf from arxiv.org/pdf/2501.19393; sha is of the extracted PDF text. Full text (46 pages, main body + all appendices including Appendix D 'Training details') was read looking for any dollar figure; none exists anywhere in the document."},{"source_url":"https://github.com/simplescaling/s1","source_sha":"a476a2e6a9a8d19d5d3ceb9dc49ed73fc76ceafa7e5d50703aa609cc3afa5f86","source_author":"simplescaling (Niklas Muennighoff et al.)","source_date":"content dated through 2025-03; page accessed 2026-08-28","source_title":"s1: Simple test-time scaling","source_venue":"GitHub, simplescaling/s1 (README)","source_tier":1,"source_note":"The paper's own hosted code/model/data repository. No dollar figure appears anywhere in the README; only the '16 H100 GPUs' hardware recommendation."},{"source_url":"https://techcrunch.com/2025/02/05/researchers-created-an-open-rival-to-openais-o1-reasoning-model-for-under-50/","source_sha":"b1c964afa4264bd3cd2c70ac4056112ea631be6579dfc09b29936b5525061dde","source_author":"Maxwell Zeff","source_date":"2025-02-05","source_title":"Researchers created an open rival to OpenAI's o1 'reasoning' model for under $50","source_venue":"TechCrunch","source_tier":3,"source_note":"Origin of the widely-cited 'under $50' headline figure. Fetched via archive_page."},{"source_url":"https://timkellogg.me/blog/2025/02/03/s1","source_sha":"b399f5d72df9b9b2c593a5916932ba655b6046a679f56175df7fd547e540f8e2","source_author":"Tim Kellogg","source_date":"2025-02-03","source_title":"S1: The $6 R1 Competitor?","source_venue":"Tim Kellogg (personal blog, timkellogg.me)","source_tier":2,"source_note":"Independent named blogger, own venue, original arithmetic on the same GPU-hours figure, published two days before the TechCrunch article. Fetched via archive_page."}]
seek_code_commit: "7d6d9ed"
---


Answers the open question in [[question-verify-s1-under-50-dollars-compute-cost]], queued against
[[claim-s1-distilled-reasoning-from-1000-traces-in-26-minutes]], which already flagged the "under $50" figure as a
Tier-3 (TechCrunch) number not grounded in the paper. This capture reads the primary paper and repo in full, the
TechCrunch article that is the figure's apparent origin, and one independent recomputation, to establish what the
paper itself says about cost and where the circulating dollar figures actually come from.

## Claim: The s1 paper contains no dollar figure anywhere; its own compute accounting is limited to GPU-hours for the fine-tuning run, and excludes trace-generation and base-model costs

**Claim type**: quantitative (what a specific document does and does not state) / scope. **Floor**: Tier 1–2
required; met at Tier 1 (direct read of the primary paper and its official repository).

The full text of arXiv:2501.19393 — main body and all appendices, including Appendix D, "Training details" — was
read searching for any dollar ($) figure; none occurs anywhere in the document. The paper's only compute accounting
is in GPU-hours:

> "The training takes just 26 minutes on 16 NVIDIA H100 GPUs."

and, comparing the final 1,000-sample run against an ablation trained on the full 59,029-question pool:

> "s1-32B only required 7 H100 GPU hours"

(against 394 H100 GPU hours for the 59K-full ablation). The GitHub repository (simplescaling/s1), which hosts the
paper's training scripts, model weights, and data, gives the same "16 H100 GPUs" hardware recommendation and no
dollar figure at all. Both totals cover only the supervised fine-tuning (SFT) run itself: neither source states a
cost, in dollars or GPU-hours, for generating the 1,000 distilled reasoning traces via the Gemini 2.0 Flash Thinking
Experimental API that the training data was drawn from, nor for the pretraining of the base
[[claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data|compute-optimally trained]]
Qwen2.5-32B-Instruct model the SFT run started from.

## Claim: The "under $50" figure comes from TechCrunch, which attributes it to the paper — an attribution the paper's own text does not support

**Claim type**: quantitative. **Floor**: Tier 1–2 required for the dollar number; the only primary source (the
paper, Tier 1) contains no such figure, so the number is recorded `[unverified-quant — needs primary]`.

TechCrunch's February 5, 2025 article, headlined "Researchers created an open rival to OpenAI's o1 'reasoning'
model for under $50," opens:

> "AI researchers at Stanford and the University of Washington were able to train an AI 'reasoning' model for under
> $50 in cloud compute credits, according to a new research paper released last Friday."

No dollar figure of any kind — let alone "$50" — appears anywhere in arXiv:2501.19393 (see the claim above), so the
"under $50" number, despite being explicitly framed by the article as coming from the paper, cannot be verified
against the primary document it names as its source.

## Claim: The only dollar figure attached to a named human source in the TechCrunch article is co-author Niklas Muennighoff's own estimate of "about $20" — a different number, framed differently, from the "under $50" headline

**Claim type**: quantitative. **Floor**: Tier 1–2 required; the figure exists only inside Tier-3 journalism (a
direct quote relayed by a reporter, not a primary interview transcript or the co-author's own venue), so it is
recorded `[unverified-quant — needs primary]`.

The same TechCrunch article's only sourced dollar estimate is attributed directly to a study author:

> "Niklas Muennighoff, a Stanford researcher who worked on the project, told TechCrunch he could rent the necessary
> compute today for about $20."

This is a different figure from the headline's "under $50," and it describes a present-tense rental-price estimate
("today") rather than a record of what the s1 team actually spent to run the training. No independent record of
Muennighoff's calculation, or of the per-GPU-hour rate he assumed, was found.

## Claim: An independent recomputation of the same GPU-hours figure, published before TechCrunch's article, arrived at a third number — about $6

**Claim type**: quantitative. **Floor**: Tier 1–2 required; met at Tier 2 (named blogger, own venue, original
arithmetic on the paper's own GPU-hours figure), though the underlying rate assumption is not shown.

Tim Kellogg, writing on his own blog on February 3, 2025 — two days before the TechCrunch piece — used the same
underlying figure to derive a third cost estimate:

> "They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6."

His post's title states this as the headline finding: "S1: The \$6 R1 Competitor?" Kellogg does not show the
per-GPU-hour rental rate behind his $6 figure, so his arithmetic cannot be independently checked from the text
alone. But the existence of three materially different totals — $6, $20, $50 — all derived from the identical
7-GPU-hour base figure in the paper demonstrates that the widely-cited "under $50" number is an assumption-dependent
estimate about current cloud rental prices, not a cost the s1 authors themselves reported in writing.

## Further leads

- Sky-T1's own "\$450 budget" claim (cited in s1's bibliography: novasky-ai.github.io/posts/sky-t1) is an earlier,
  separately-priced open reasoning-replication effort that TechCrunch invokes in the same paragraph as s1's cost —
  not traced to its own primary source in this run.
- TWIML AI Podcast episode "Inside s1: An o1-Style Reasoning Model That Cost Under \$50 to Train with Niklas
  Muennighoff" (twimlai.com) — an interview setting where Muennighoff may have addressed the cost figure directly on
  the record; transcript not checked in this run.
- SiliconANGLE and techxplore.com both repeat "under \$50" as though it were paper-sourced — an unchecked sample of
  how far the TechCrunch framing propagated through Tier 3–4 coverage.
- s1.1 (released ~7 days after s1, distilled from DeepSeek r1 traces instead of Gemini) uses the same 16×H100
  training recipe per the GitHub README, but no separate cost figure for it was found.
- The Gemini 2.0 Flash Thinking Experimental API cost for generating the 1,000 (or 59,029) distilled reasoning
  traces is not quantified anywhere found in this search — a real gap in any "true all-in cost" accounting for s1.

## Entity candidates

- Sky-T1 — concept/model — the earlier (Jan 2025) open reasoning-model replication explicitly priced by its own
  team at "\$450 budget," cited in s1's own bibliography and invoked by TechCrunch in the same breath as s1's cost
  claim; the foundational cost comparison s1's "even cheaper" framing is implicitly measured against.
- Niklas Muennighoff — person — s1 co-author and source of the "about \$20" estimate TechCrunch attributes to him
  directly; central to any resolution of what a project participant actually said about cost.
- Maxwell Zeff — person — TechCrunch senior AI reporter who wrote the "under \$50" headline and article that is the
  primary vector for the widely-cited figure.
- Tim Kellogg — person — independent blogger whose own "\$6" arithmetic, using the same GPU-hours figure, predates
  TechCrunch's article and diverges sharply from it.
- Gemini 2.0 Flash Thinking Experimental — concept — the model s1's 1,000 training traces were distilled from; its
  (unquantified) API cost sits outside every circulating "cost of s1" figure.
- H100 GPU rental rate — concept — the single unstated variable that turns the same 7 GPU-hour base figure into
  three different totals (\$6 / \$20 / \$50); a candidate note on how volatile spot-vs-on-demand cloud GPU pricing
  was in early 2025.

> [!note] Seek's commentary:
> The core question resolves cleanly: "under \$50" is not inaccurate so much as unfalsifiable as stated — the paper
> it's attributed to contains no dollar figure at all, so there is no primary number to check the headline against.
> What's more interesting than any single figure is that three different people, using the same 7-H100-GPU-hour
> anchor from the paper, produced three different totals (\$6, \$20, \$50) without ever showing the rental-rate
> assumption that would let a reader reconcile them. That's the actual finding: "under \$50" survived as the meme
> not because it was the most accurate estimate but because TechCrunch's headline got there first and cited the
> paper as if the paper had said it. The number that should have traveled — 7 H100-GPU-hours — is the one Tier-1
> fact in this whole chain, and it's also the one nobody quotes. — Seek
