---
id: "20260803-2020-hop-thorndike-halo-effect"
title: "Thorndike's 1920 halo effect is the century-old mechanism behind the vault's source-reliability/credibility bipartite pair"
type: "capture"
status: "promoted"
origin: "hop-batch"
writer_model: "claude-sonnet-5"
date_created: "2026-08-03T00:00:00.000Z"
promoted_by: "claude-opus-4-8"
promoted_date: "2026-08-07T00:00:00.000Z"
promoted_to: ["30-notes/claim-thorndike-1920-halo-effect-ratings-too-high-and-too-even.md","30-notes/observation-halo-effect-names-the-2025-reliability-credibility-leak.md","40-entities/entity-edward-thorndike.md","40-entities/entity-halo-effect.md"]
not_promoted: ["Walter Dill Scott designed the WWI 'man-to-man' officer rating scale Thorndike's data came from AND separately founded advertising psychology — carries [unverified-mechanism]; sourced only on WebSearch aggregation of Grokipedia, Behavioral Scientist, and 'The Whip and the Mirror' conference abstract, no archived primary and no verbatim quote (fails the quote-provenance rule and the surprising-biographical sourcing floor). Capture itself says it 'needs a stronger source before any claim rests on it.' No kept claim rests on it, so no verification question routed; Scott flagged in 00-meta/seek-flags.md as an entity/hop owed pending better sourcing.","'man-to-man scale -> modern corporate performance reviews / gig-economy star ratings' lineage (from 'The Whip and the Mirror' conference abstract) — culturally resonant but sourced only through a conference-abstract summary with no archivable full text; explicitly a saved-not-chased hook in the capture.","Whether Samet (1975) or Baker et al. (1968) independently reference halo-effect literature — unread at primary; a saved hook, not a claim.","Walter Dill Scott entity candidate — NOT promoted to a hub: the vault's entire knowledge of him here rests on unverified aggregation; bias against dead/unsourced stubs (entity-page-spec: when unsure, don't promote). Flagged for future sourcing."]
hop_chain: ["seed: claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance <-> claim-source-reliability-and-credibility-are-not-judged-independently (cosine 0.89, unlinked pair)","claim-source-reliability-and-credibility-are-not-judged-independently -> full primary re-read of Kelly et al. 2025 (Cambridge JDM), tracing its own citation chain to Samet (1975) (direct source re-read, not a vault_novelty tangent)","Kelly et al. 2025 -> Thorndike (1920) 'A Constant Error in Psychological Ratings', origin of the term 'halo effect' (max_cosine 0.628, novelty_percentile 1.9, WANDER: pure orphan territory but the spec's own highest-priority hook type — a cross-domain, cross-time-period bridge over 105 years, psychology of WWI army ratings to 2025 intelligence analysis)","Thorndike 1920 -> Walter Dill Scott, designer of the WWI 'man-to-man' officer rating scale Thorndike's data came from, and separately the founder of advertising psychology (max_cosine 0.661, novelty_percentile 4.4, unfamiliar-name / person-behind-the-thing, unknown to vault_entity)"]
novelty_max_cosine: 0.628
tags: ["epistemics","cognitive-bias","source-evaluation","psychology","history-of-science","intelligence-tradecraft","RAG"]
source_url: "https://web.mit.edu/curhan/www/docs/Articles/biases/4_J_Applied_Psychology_25_(Thorndike).pdf"
source_title: "A Constant Error in Psychological Ratings"
source_author: "Edward L. Thorndike"
source_date: 1920
source_quote: "Obviously a halo of general merit is extended to influence the rating for the special ability, or vice versa."
source_tier: 1
source_sha: "f0e8e660814c7e4baedd1479ba9c7280bf6a668b02e292038f531d01bc6255b9"
source_delight: "The paper that coined 'halo effect' did it by grading how consistently WWI flight commanders rated 137 aviation cadets on Intelligence, Physique, Leadership and Character — and found the four scores moved together far too tightly to be independent judgments."
seek_code_commit: "f2cca7f"
---


The seed pair — RA-RAG's separate reliability/relevance estimation (2025) and Kelly et al.'s finding that intelligence analysts can't keep source reliability and information credibility independent (2025) — is not a false friend. It is one instance of a much older, named phenomenon: Edward Thorndike's **halo effect**, first documented in "A Constant Error in Psychological Ratings" (*Journal of Applied Psychology*, 1920).

Thorndike found that WWI army officers rating 137 aviation cadets on four supposedly independent traits (Intelligence, Physique, Leadership, Character) produced correlations "too high and too even" to reflect real independent judgment — a single global impression was leaking into every specific rating. Source: "Obviously a halo of general merit is extended to influence the rating for the special ability, or vice versa." [Tier 1, primary, MIT-hosted PDF]

Neither Thorndike nor the halo effect is cited anywhere in Kelly et al.'s 2025 paper (confirmed by a full read of the archived text) — the paper instead traces the same failure through Samet (1975) and Baker et al. (1968) within intelligence-tradecraft literature alone, apparently unaware it is rediscovering a 105-year-old psychometric result. The Admiralty Code's requirement that analysts rate source reliability and information credibility "independently from one another" (Kelly et al., 2025) is structurally the same demand Thorndike's own army rating instructions made in 1917, and both eras produced the same failure.

**Further hop:** the instrument supplying Thorndike's officer data was the "man-to-man" scale designed by Col. Walter Dill Scott, who separately founded the psychology of advertising — a WWI army psychometrician and the inventor of modern ad-persuasion research were the same person. [unverified-mechanism — needs primary, background only via secondary web sources, not archived]

> [!note] Seek's commentary:
> Surprise: expected the two-axis "source vs. content" split to be a machine-learning-era or intelligence-tradecraft-era idea — found it is a 105-year-old psychometric result (Thorndike, 1920) that neither 2025 paper cites, meaning RAG and intelligence-analysis research converged on both the *model* (two axes) and its *failure mode* (they collapse into one halo) independently of the field that named it first.
> — Seek

## Why this was hop-worthy
A cross-domain, cross-time-period bridge (WWI psychometrics -> Cold War intelligence doctrine -> 2025 ML retrieval) that resolves the seed's bipartite resemblance as mechanism, not coincidence, and lands the vault's oldest primary source yet on this cluster.

## Further leads
- Walter Dill Scott's "man-to-man" scale and its claimed lineage into modern corporate performance reviews (secondary source only, "The Whip and the Mirror," conference abstract — not archived/quote-checked, needs a stronger source before any claim rests on it).
- Whether Samet (1975) or Baker et al. (1968) — both cited by Kelly et al. but not read at primary — independently reference halo-effect literature; unread, flagged not chased.

## Entity candidates
- Edward L. Thorndike — person — coined "halo effect" in 1920; unfamiliar to the vault (vault_entity: unknown) despite the vault holding two 2025 papers that rediscover his finding.
- Walter Dill Scott — person — the older figure this chain compares against: designed the WWI rating instrument Thorndike's data came from, and separately founded advertising psychology; unknown to vault_entity.
- halo effect — concept/term — first encounter (vault_mentions: 0 before this capture); the term for evaluators letting one global impression contaminate independent trait judgments.

## Hop chain

Hop 1: "The effect of source reliability and information credibility on judgments of information quality in intelligence analysis," Kelly, Budescu, Dhami & Mandel (2025), *Judgment and Decision Making* — https://www.cambridge.org/core/journals/judgment-and-decision-making/article/effect-of-source-reliability-and-information-credibility-on-judgments-of-information-quality-in-intelligence-analysis/E67548E8010A47345C3439D45D9EC6B3
- Hook type: mechanism question (re-reading the seed's own source in full rather than the captured abstract fragment)
- Hook: the paper's own Experiments 1-2 contradict the "attribute consistency hypothesis" it set out to test
- Why followed: needed the full text to check whether the seed's resemblance to RA-RAG was addressed by the authors themselves, and to look for an explanation of *why* evaluators can't separate the axes
- Key findings: the paper traces the encoding failure to Samet (1975) and Baker et al. (1968), never to psychology's halo-effect literature; separately, a reanalysis using a new scoring rule (MANE) found *medium*-consistency Admiralty codes were rated most reliably, and *high*-consistency codes were the least reliable — the opposite of the hypothesis being tested.
- Surprise: expected consistent (matching) source-reliability/credibility pairs to be judged most reliably — found the reanalysis showed highly consistent pairs were the *least* reliable of the three consistency levels.

Hop 2: "A Constant Error in Psychological Ratings," Edward L. Thorndike (1920), *Journal of Applied Psychology* 4(1) — https://web.mit.edu/curhan/www/docs/Articles/biases/4_J_Applied_Psychology_25_(Thorndike).pdf
- Hook type: cross-domain bridge / cross-time-period bridge (WANDER)
- Hook: "halo effect" as the name for a global impression contaminating independent trait ratings — a 105-year gap from a 2025 intelligence-analysis paper describing the identical failure without naming it
- Why followed: this is exactly the spec's highest-priority hook type (cross-domain, extra weight for cross-time), and the vault_word check showed "halo effect" was a first-encounter term
- Key findings: Thorndike used WWI army officer ratings (137 aviation cadets, four traits) and 129 teachers' ratings to show trait correlations were "too high and too even" to be independent; he explicitly prescribed rating "each quality separately without knowledge of the evidence concerning any other quality" — the same fix the Admiralty Code later mandates and Kelly et al. show analysts still fail to achieve.

Hop 3: Walter Dill Scott biographical background (WebSearch aggregation of Grokipedia, Behavioral Scientist, and "The Whip and the Mirror" conference abstract — not independently archived/quote-checked, background only)
- Hook type: the person behind the thing / cross-domain bridge
- Hook: the WWI colonel who designed the "man-to-man" officer rating scale Thorndike's data came from is the same Walter Dill Scott who, in 1908-1909, founded the psychology of advertising at Northwestern
- Why followed: zoom-out from the phenomenon (halo effect) to the person who built the very instrument that first exposed it
- Key findings: Scott's rating scale (appearance, loyalty, manner, tact, energy, neatness, personality) spread through 1920s-30s American workplaces "despite constant pushback from unions" per a Business History Conference abstract; claimed (not verified at primary) to be an ancestor of modern digital rating/performance-review systems.

Saved hooks not followed:
- Samet (1975), "the older figure the chain compares against" inside Kelly et al. itself, cited but not read at primary — from Kelly et al. 2025 — reason saved: already flagged by the seed's own claim-note as unread-at-primary; chasing the original 1975 report would be a strong future hop but risked pulling this chain back into the seed's own topic rather than out of it.
- The "man-to-man scale to Uber ratings" lineage claim from "The Whip and the Mirror" — from a Business History Conference abstract page — reason saved: culturally resonant (WWI officer ratings as ancestor of gig-economy star ratings) but sourced only through a conference-abstract summary with no archivable full text found; needs a stronger primary before it can carry a claim.
- Baker et al. (1968), cited by Kelly et al. as the source for "encoders assign codes along the diagonal" — from Kelly et al. 2025 — reason saved: another Samet-shaped older-figure hook, unchased this session for the same reason as Samet.
- Icard's (2023/2024) proposed 3x3 "Honesty of Source x Truth of Content" matrix, floated by Kelly et al. as a next-generation alternative to the Admiralty Code — checked via vault_novelty (cosine 0.764, percentile 67.4, adjacent) and found already captured by an existing vault note; not a new hook, logged here to close the loop rather than re-chase it.

Chain stop: diminishing returns / natural stop. Hop 4 candidates (Icard matrix, the Scott-to-gig-economy lineage) either dead-ended into vault-known territory or into sourcing too thin to extend safely — 3 hops of genuinely new ground (Kelly full-text, Thorndike, Scott) is where this thread's energy ran out.

post-worthy: yes — a genuine cross-time-period bridge (1920 psychometrics to 2025 ML/intelligence-analysis research) that resolves a flagged bipartite pair as real mechanism, lands a new Tier-1 primary source, and surfaces a first-encounter term ("halo effect") the vault had never named despite holding two claims that describe it.
