s1 reproduced o1-preview-level math reasoning by fine-tuning Qwen2.5-32B on 1,000 distilled traces in 26 minutes
The s1 model (Muennighoff et al., "s1: Simple test-time scaling," arXiv:2501.19393, 2025-01-31) was built by "supervised fine-tuning the Qwen2.5-32B-Instruct language model" on "1,000 carefully curated questions paired with reasoning traces and answers distilled from Gemini Thinking Experimental," "requiring just 26 minutes of training on 16 H100 GPUs." The resulting model "exceeds o1-preview on competition math questions by up to 27%."
The result is evidence that a reasoning capability can transfer from a very small set of visible reasoning traces: once the chains of thought are exposed, reproducing much of the behaviour is a short, cheap supervised-fine-tuning run on an existing open base model rather than a from-scratch training effort. This is the empirical counter to the premise behind hiding the trace in claim-openai-hid-o1-raw-chain-of-thought-partly-for-competitive-advantage, and it makes reasoning look like the inverse of a tacit moat — see claim-a-chain-of-thought-trace-is-codified-so-it-cannot-form-a-tacit-moat. The mechanism it exploits — spending sequential inference steps to reach an answer — is the serial test-time compute paradigm, and it is consistent with the broader finding that test-time compute can substitute for parameters.
The flagged cost figure
The much-repeated "under $50 in compute" headline for s1 comes from TechCrunch
(a Tier-3 source), not the paper, and is carried here under
[unverified-quant — needs primary]. The paper grounds only the underlying
run — 16 H100s for 26 minutes — not a dollar figure. The verification (find a
primary cost figure or reproduce the arithmetic) is queued at
question-verify-s1-under-50-dollars-compute-cost. Because the note's most
quoted number rests on a soft source, it is held at seedling.
Source
“requiring just 26 minutes of training on 16 H100 GPUs”
claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-11-hop-cot-not-a-tacit-moat.md, 2026-07-12 · raw markdown