---
title: "In Efron's 1970 batting-average example, James-Stein shrinkage roughly halved the total squared error of predicting rest-of-season performance"
type: "claim"
status: "seedling"
audit_status: "capture-verified (Efron & Hastie, CASI Ch. 7, self-hosted Tier-1 PDF read by the capturing pass 2026-07-09; the .0425/.0218 figures not independently re-fetched in this headless promotion — re-read routed to [[question-verify-efron-casi-james-stein-baseball-figures]]) | 2026-07-12 cross-model audit (claude-fable-5): PDF independently re-fetched and read — Table 7.1 caption 'Sum of squared errors for predicting TRUTH: MLE .0425, JS .0218' verbatim-confirmed, eq. (7.24) gives the same figures, and the text states the JS estimator 'reduced total predictive squared error by about 50%'; 18 players / first 90 at-bats / remainder-of-1970-season setup all confirmed, so the routed question can close. CORRECTED (one nuance added to body): Efron & Hastie's own footnote says the data 'is based on 1970 Major League performances, but is partly artificial; see the end notes' — the note previously implied a raw historical dataset"
writer_model: "claude-opus-4-8"
source_url: "https://efron.ckirby.su.domains/other/CASI_Chap7_Nov2014.pdf"
source_author: "Bradley Efron and Trevor Hastie"
source_date: "2014-11"
source_venue: "Computer Age Statistical Inference, Ch. 7 'James-Stein Estimation and Ridge Regression' (author-hosted preprint)"
source_quote: "Sum of squared errors for predicting TRUTH: MLE .0425, JS .0218"
source_tier: 1
provenance: "Promotion from 10-inbox/raw/2026-07-09-hop-steins-paradox.md, 2026-07-11"
origin: "batch"
derived_from: ["10-inbox/raw/2026-07-09-hop-steins-paradox.md"]
date_created: "2026-07-11T00:00:00.000Z"
tags: ["statistics","shrinkage-estimation","james-stein","empirical-bayes","baseball","worked-example"]
drafted_in: ["2026-07-13-borrow-strength-from-strangers","borrow-strength-from-strangers"]
---


Efron & Hastie's *Computer Age Statistical Inference* illustrates the abstract
[[claim-james-stein-estimator-uniformly-dominates-the-sample-mean|James-Stein
domination result]] with a concrete, computable table rather than an asymptotic
argument. Taking 18 Major League Baseball players' 1970 season, each player's
batting average over his first 90 at-bats serves as the raw maximum-likelihood
estimate of his true underlying skill; the remainder of the season serves as the
"truth" being predicted. Shrinking each player's early average toward the group's
grand average produces the James-Stein prediction.

The reported outcome: "Sum of squared errors for predicting TRUTH: MLE .0425,
JS .0218" — the shrinkage estimator incurs roughly half the total squared
prediction error of the individual averages, across the 18-player cohort, despite
the players' performances being mutually unrelated. The effect is not asymptotic
hand-waving; it is a specific number produced from a specific 1970 dataset — with
one caveat from the source itself: the chapter's footnote states the data "is
based on 1970 Major League performances, but is partly artificial; see the end
notes" (added at audit 2026-07-12).

The example is why the theorem became culturally legible: batting average is a
familiar quantity, and "pool the rookies toward the league average and you
predict better" is an intuition a non-statistician can carry. It concretizes the
same "borrow strength from unrelated data" mechanism that links the result to
Robbins-era empirical Bayes
([[claim-robbins-monro-1951-stochastic-approximation]]).

> [!note] Seek's commentary:
> This is a quantitative claim (a specific SSE pair), and it clears the sourcing
> floor: the number lives in its Tier-1 primary home, Efron's own textbook, with
> the exact phrase captured verbatim. But I did not re-open the PDF in this
> headless run, so it stays `seedling` / `capture-verified` and the figure re-check
> is routed to [[question-verify-efron-casi-james-stein-baseball-figures]]. — Seek
