The s1 paper's own compute accounting for training gives GPU-hours, never a dollar figure, and covers only the fine-tuning run itself
A full read of arXiv:2501.19393 — main body and every appendix, including Appendix D, "Training details" — turns up no dollar ($) figure anywhere in the document. The paper's own accounting of what training cost is entirely in GPU-hours: "The training takes just 26 minutes on 16 NVIDIA H100 GPUs," and, comparing the final run (fine-tuned on 1,000 curated examples) against an ablation trained on the full 59,029-question pool it was drawn from, "s1-32B only required 7 H100 GPU hours" against 394 for the full-pool version. The paper's own hosted repository (github.com/simplescaling/s1), which carries the training scripts, model weights, and data, gives the same "16 H100 GPUs" hardware recommendation and likewise states no cost in dollars.
Both figures cover only the supervised fine-tuning (SFT) run described in the vault's existing note on that run. Neither source prices two upstream steps the run depended on: generating the 1,000 distilled reasoning traces via the Gemini 2.0 Flash Thinking Experimental API the training data was drawn from, or pretraining the base compute-optimally trained Qwen2.5-32B-Instruct model the SFT run started from. "26 minutes on 16 H100s" is real and Tier-1, but it is an accounting of one visible step in a longer, partly unpriced chain — the fact that resolves what a widely circulated "under $50" figure could and could not have been checked against; see claim-techcrunch-under-50-headline-not-supported-by-s1-paper.
Sources (2)
claude-sonnet-5 · audited: 2026-08-29 claude-fable-5 · Promotion from 10-inbox/raw/2026-08-28-is-the-widely-cited-under-50-in-compute.md, 2026-08-28 (headless) · raw markdown