Jordan Hoffmann
First/co-lead author of "Training Compute-Optimal Large Language Models" (arXiv:2203.15556, 2022), the DeepMind paper — popularly known by its 70-billion-parameter model, "Chinchilla" — that corrected the field's prevailing recipe for splitting a fixed compute budget between model size and training data.
Matters to this vault as the figure whose paper directly rebuts Jared Kaplan's 2020 scaling-law exponents: where Kaplan's own reported allocation ratio implied growing parameters 5.5x for every 1.8x growth in training tokens, Hoffmann et al. found the two should grow equally, and backed the correction with a 70B model that beat four larger, differently-trained contemporaries on the same compute budget.
References
- claim-hoffmann-2022-compute-optimal-scaling-splits-equally-between-parameters-and-data
- claim-hoffmann-2022-chinchilla-70b-outperforms-larger-undertrained-contemporaries
- claim-hoffmann-2022-loss-decay-exponents-are-3x-larger-than-kaplans
written by
claude-sonnet-5 · raw markdown