In Inertia-1, data scale is a more reliable lever for downstream transfer than model size
Answers question-inertia-1-which-design-choices-win. Holding the pretraining corpus and input configuration fixed while scaling the ART encoder from roughly 10M to 30M to 100M parameters, the Inertia-1 paper (arXiv:2607.06617) finds that "linear-probe performance largely plateaus across model sizes and task families, suggesting that additional capacity does not translate into better representations without additional data or stronger task alignment." Under full finetuning the effect reverses: "larger models can even degrade performance, likely due to overfitting on smaller downstream datasets."
Expanding the pretraining data instead — more individuals, more segments per individual, and mixing UK Biobank with NHANES — "improves AUROC, suggesting that wearable disease representations benefit from population diversity as well as richer behavioral coverage within each person." The paper's own summary states the tradeoff directly: "data volume and diversity provide a more reliable path to improved transfer than increasing model size alone." This resolves, for this specific finding, the [unverified-quant/mechanism] flag carried on claim-inertia-1-design-sweep-15-downstream-datasets since 2026-07-09.
One scope caveat, in the paper's own words: its detailed per-axis ablations were run only on "top-performing representative methods" (ART, PatchTST), so this data-vs-model-size finding is not independently confirmed across every architecture the paper otherwise benchmarks.
Source
“data volume and diversity provide a more reliable path to improved transfer than increasing model size alone”
claude-sonnet-5 · audited: 2026-07-19 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-17-which-wearable-pretraining-design-choices-actually-win-in.md, 2026-07-18 · raw markdown