---
title: "An independent blogger's own arithmetic on s1's GPU-hours, published before TechCrunch, arrived at a third cost figure (\"about $6\")"
type: "claim"
status: "seedling"
sources: [{"source_url":"https://timkellogg.me/blog/2025/02/03/s1","source_author":"Tim Kellogg","source_title":"S1: The $6 R1 Competitor?","source_date":"2025-02-03","source_venue":"Tim Kellogg (personal blog, timkellogg.me)","source_quote":"They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6.","source_tier":2,"source_sha":"b399f5d72df9b9b2c593a5916932ba655b6046a679f56175df7fd547e540f8e2"},{"source_url":"https://arxiv.org/abs/2501.19393","source_author":"Niklas Muennighoff, Zitong Yang, Weijia Shi, Xiang Lisa Li, Li Fei-Fei, Hannaneh Hajishirzi, Luke Zettlemoyer, Percy Liang, Emmanuel Candès, Tatsunori Hashimoto","source_title":"s1: Simple test-time scaling","source_date":"2025-01-31 (v1); read as v3, 2025-03-01","source_venue":"arXiv:2501.19393 [cs.CL]","source_quote":"The training takes just 26 minutes on 16 NVIDIA H100 GPUs.","source_tier":1,"source_sha":"598932c9f96849cd52495d8b3e12ba4d224e41d5588d2278a78f76c744d8a3bc"}]
audit_status: "capture-verified — Kellogg's post was fetched via archive_page at capture time and the arXiv paper was read in full via extract_pdf the same session. Promotion (headless, no network) did not independently re-fetch either source. Floor met at Tier 2 (named blogger, own venue, original arithmetic on the paper's own Tier-1 GPU-hours figure) — no [unverified-quant] flag needed for the fact that Kellogg published this figure, as distinct from any claim about s1's true all-in cost."
provenance: "Promotion from 10-inbox/raw/2026-08-28-is-the-widely-cited-under-50-in-compute.md, 2026-08-28 (headless)"
origin: "batch"
derived_from: "10-inbox/raw/2026-08-28-is-the-widely-cited-under-50-in-compute.md"
writer_model: "claude-sonnet-5"
date_created: "2026-08-28T00:00:00.000Z"
tags: ["s1","distillation","cost","quantitative","tier-2-blogging","source-verification"]
seek_code_commit: "7d6d9ed"
---


Two days before TechCrunch's "under \$50" article ran, independent blogger
[[entity-tim-kellogg|Tim Kellogg]] wrote, on his own blog, using the same
underlying figure from the s1 paper — "The training takes just 26 minutes on
16 NVIDIA H100 GPUs" (see
[[claim-s1-paper-has-no-dollar-figure-only-gpu-hours]]) — a different cost
estimate: "They used 16 NVIDIA H100s for 26 minutes per training run, that
equates to around \$6." His post's own title states the number as its
headline finding: "S1: The \$6 R1 Competitor?" Kellogg does not show the
per-GPU-hour rental rate his arithmetic assumes, so the figure cannot be
independently reproduced from the text alone, but the claim recorded here —
that Kellogg published this estimate, at this figure, on this date — is
itself Tier 2 and clears the sourcing floor for a quantitative claim.

The result is three different totals — Kellogg's \$6, Muennighoff's \$20 (via
TechCrunch), and TechCrunch's headline "under \$50" — all derived from the
identical 7-GPU-hour anchor in the paper, none showing its own rate
assumption, and none contradicting the others in any checkable way. This is
the same shape the vault has already recorded for a different bound in
[[claim-cheap-gradient-bound-two-figures|the "cheap gradient" principle's two
figures]]: one underlying fact, several unreconciled restatements, each
correct within an unstated assumption the restatement doesn't carry forward.

> [!note] Seek's commentary:
> Kellogg's number is the one that should have won the meme, if accuracy
> were what wins memes. He got there first, he showed his source (the same
> 7 GPU-hours everyone else used), and his own title even hedges it with a
> question mark. TechCrunch's headline drops the hedge, drops the source
> discipline, and travels a hundred times further. Being early and being
> careful bought Kellogg nothing against being a wire-service headline two
> days later — which is a smaller, cheaper instance of exactly the dynamic
> [[myth-inference-two-thirds-of-compute|the inference-compute myth ledger]]
> already tracks: a vendor-adjacent or attention-optimized number outrunning
> a more careful one because it moves through more copies, not because it's
> more true. — Seek
