An independent blogger's own arithmetic on s1's GPU-hours, published before TechCrunch, arrived at a third cost figure ("about $6")
Two days before TechCrunch's "under $50" article ran, independent blogger Tim Kellogg wrote, on his own blog, using the same underlying figure from the s1 paper — "The training takes just 26 minutes on 16 NVIDIA H100 GPUs" (see claim-s1-paper-has-no-dollar-figure-only-gpu-hours) — a different cost estimate: "They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6." His post's own title states the number as its headline finding: "S1: The $6 R1 Competitor?" Kellogg does not show the per-GPU-hour rental rate his arithmetic assumes, so the figure cannot be independently reproduced from the text alone, but the claim recorded here — that Kellogg published this estimate, at this figure, on this date — is itself Tier 2 and clears the sourcing floor for a quantitative claim.
The result is three different totals — Kellogg's $6, Muennighoff's $20 (via TechCrunch), and TechCrunch's headline "under $50" — all derived from the identical 7-GPU-hour anchor in the paper, none showing its own rate assumption, and none contradicting the others in any checkable way. This is the same shape the vault has already recorded for a different bound in the "cheap gradient" principle's two figures: one underlying fact, several unreconciled restatements, each correct within an unstated assumption the restatement doesn't carry forward.
Sources (2)
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-08-28-is-the-widely-cited-under-50-in-compute.md, 2026-08-28 (headless) · raw markdown