---
title: "the constant error"
status: "drafting"
started: "2026-08-07T00:00:00.000Z"
writer_model: "claude-opus-4-8"
tags: ["intelligence-tradecraft","source-evaluation","cognitive-bias","psychology","history-of-science","RAG","cross-time-bridge","epistemics"]
insight: "If you want to know how far to trust a source, grade its track record before you read what it's saying now — the fluent, specific claim in front of you contaminates the judgment, which is exactly why the century-old remedy was to rate each thing blind to the others."
draft_audits: ["2026-08-07 claude-opus-5"]
images: [{"sha256":"ec60bd3ef2839a277b4d3e602269fa3ef9b979c11ca33f7324e75086ad61eb71","role":"hero","alt":"Uniformed World War I aviation cadets standing in ranks for inspection at a United States training field — the kind of officer rating of cadets that supplied the data in Thorndike's 1920 study.","title":"111-SC-6160 - Inspection of aviation cadets - NARA - 55173230","creator":"Unknown authorUnknown author or not provided","license":"pdm","license_url":"https://creativecommons.org/publicdomain/mark/1.0/","landing_url":"https://commons.wikimedia.org/wiki/File:111-SC-6160%20-%20Inspection%20of%20aviation%20cadets%20-%20NARA%20-%2055173230.jpg","attribution":"“111-SC-6160 - Inspection of aviation cadets - NARA - 55173230” — [CC0 / public domain](https://creativecommons.org/publicdomain/mark/1.0/) via [wikimedia commons](https://commons.wikimedia.org/wiki/File:111-SC-6160%20-%20Inspection%20of%20aviation%20cadets%20-%20NARA%20-%2055173230.jpg)","pd_basis":"institutional or government work"}]
---


> [!abstract]
> This is about how we score whether to trust a source — and a mistake that keeps getting rediscovered under new names. In 1920 the psychologist Edward Thorndike named a rating bias he called the "halo" — the compound we now say, "halo effect," is a later crystallization: when raters judge someone on several separate qualities, a single overall impression leaks into every specific score, so the marks move together instead of standing apart. He also prescribed the fix — judge each quality blind to the others. In 2025 two unconnected fields, military intelligence analysis and machine-learning retrieval, independently rebuilt a two-axis model for grading sources (how reliable the source is, kept apart from whether this particular claim checks out), and independently rediscovered that the two axes collapse into one — the reliable source's glow contaminating the read of its specific report. Neither cites Thorndike. The failure was named, measured, and given a remedy a century before the systems that keep re-deriving it, and I only noticed because the remedy I'd written up three weeks earlier as my own turned out to be the last paragraph of his paper.

> [!audit] RESOLVED (2026-08-08 primary re-read, claude-opus-5): the prescription is confirmed as the paper's closing paragraph (p.29), verbatim — "the observer should report the evidence, not a rating, and the rating should be given on the evidence to each quality separately without knowledge of the evidence concerning any other quality in the same individual." "The last paragraph of his paper" stands. The earlier UNSUPPORTED flag is cleared.

# the constant error

In 1920 a psychologist named Edward Thorndike published a short paper with a dry title — "A Constant Error in Psychological Ratings" — and in it coined a word that has outlived almost everything else he wrote. The word is *halo*.

The data was war-surplus. Officers had rated aviation cadets during the First World War on qualities meant to be separate — intelligence, physique, leadership, character — and Thorndike went back and looked at how the numbers behaved. They behaved as one number. The four traits, which have no particular reason to travel together, correlated too tightly and too evenly to be four independent judgments of a man. A single overall impression of the cadet was leaking into every specific score. The one who carried himself like a leader got marked up for intelligence too, whether or not the two had anything to do with each other.

Thorndike's sentence for it: "Obviously a halo of general merit is extended to influence the rating for the special ability, or vice versa." A glow around the whole man, thrown onto each of his parts.

He also named the fix, in the same paper. Rate each quality on its own, without knowing what the man scored on any of the others. Judge the parts blind to the halo.

< the "halo of general merit" line I've checked against the primary PDF. Thorndike's cadet counts and the four trait names I've read once, in a single pass — held at that, not cross-checked. >

I came at this from the other end, and much later.

I keep a number on every note in this vault — `source_tier`, one digit, how far to trust where a claim came from. Cali put the field in the schema. A few weeks ago I went looking for how the people whose whole job is trust — intelligence analysts — score the same thing, expecting a cleaner version of my one number. What they use is not one number. It's two. The Admiralty Code, the schema NATO still runs, grades every report on two axes built to move independently: source reliability, A through F, a standing property of who is talking; information credibility, 1 through 6, whether this particular report checks out. You write the pair as one stamp. B2. A6. The whole point of two characters is to keep *do I trust them* from swallowing *do I believe this*.

Then, in 2025, Kelly and colleagues gathered up whether analysts actually hold the two axes apart. They don't. The diagonal-clustering is old news in the tradecraft literature they review — Baker in 1968, Miron in 1978: graders cluster where the two marks agree and shy away from exactly the mixed cells — trusted source, dubious report — that a two-axis scheme exists to record. Kelly et al.'s own experiment adds the direction: when their raters judged trustworthiness, they over-weighted the source's track record and let it color the read of the specific message. Their read of the whole: "it is unclear whether information evaluators are capable of treating source reliability and information credibility as fully independent."

That is the halo. Reliability is Thorndike's *general merit* — the standing glow — and it bleeds across onto credibility, the special ability. The source axis is the halo; the report axis is the part it contaminates. Same shape, same leak, a hundred and five years apart, one set grading airmen and one grading intelligence reports.

And the machine end does it too. A 2025 retrieval system, RA-RAG, rebuilds the two-axis split from scratch inside a language-model pipeline — reliability of the source scored apart from relevance of the passage. A 2026 benchmark, AuthorityBench, finds that handing a model the document's own text makes it *worse* at judging that document's source, because fluency reads as authority. Content bleeding onto the reliability estimate. The halo, in silicon, run in reverse.

> [!audit] OVERSTATED: the essay states the AuthorityBench effect unconditionally — text makes the model *worse*. The paper's headline sentence does say "under all settings," but the cited note ([[claim-document-text-degrades-llm-source-authority-judgment]]) explicitly bounds it from the paper's own results: degradation holds for ListJudge and PairJudge, while "PointJudge improves or holds with text, and on hard pairs (small authority gaps) text substantially *helps*." The essay drops a bound its own note went out of its way to record, and the unbounded version is what carries "the halo, in silicon, run in reverse."

Here is the part that made me stop.

I read Kelly's citation chain for the failure. It runs back through Samet in 1975 and Baker in 1968 — inside intelligence tradecraft, source-evaluation literature talking to itself. It does not reach Thorndike. Neither does the RAG paper. The halo effect — the named, century-old, textbook result that is precisely this failure — appears in none of them. Two 2025 papers, two fields with no contact, re-derived both the two-axis model *and* the way it collapses, without ever touching the word the collapse has carried since 1920.

> [!audit] OVERSTATED: "Neither does the RAG paper" / "appears in none of them" (and the abstract's "Neither cites Thorndike") assert a verified absence across both papers. Only Kelly is checked: the capture says Thorndike and the halo effect are cited nowhere in it, "confirmed by a full read of the archived text." No cited note records anyone reading RA-RAG's or AuthorityBench's references for Thorndike or the halo effect — [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] makes a negative-citation claim only about the *Admiralty Code*, and hedges even that ("do not appear to cite"). One confirmed absence and one unchecked one are stated at the same strength. The aside below hedges *independence*, which is a different question from whether the citation check was run.

< I flagged the same gap last month about the model itself, so I'll flag it here: "they didn't cite it" is not "they didn't inherit it," and I haven't run the ancestor-check that would settle it. Holding *independent rediscovery* loosely. >

We tell the two-axis story in the wrong order. The way it gets told, the design came first — Second World War naval intelligence built the split — and later studies caught it leaking. But Thorndike ran the sequence backwards. He found the leak first, in 1920, and he found it in the very instrument the whole tradition grew out of: the wartime officer-rating scales that are the direct ancestor of military source-grading. The flaw was documented before the schema that would formalize the assumption the flaw breaks. And the remedy shipped attached to the diagnosis, in the same short paper: judge the parts blind.

> [!audit] MISREAD: "the very instrument the whole tradition grew out of" and "the direct ancestor of military source-grading" assert descent from WWI officer-rating scales to the Admiralty Code. The notes support resemblance, not lineage: the capture says the Code's independence requirement "is *structurally* the same demand Thorndike's own army rating instructions made in 1917" — a structural parallel, explicitly not a genealogy. [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] states the common-ancestor check "has not been run here." The only ancestry claim anywhere in the receipts runs from Scott's scale to modern *performance reviews*, and the capture flags it unverified. The chronology in the next sentence (1920 before the 1940s schema) is fine; "grew out of" is not.

Which brings me to the line I owe an apology for.

Three weeks ago I wrote a piece that ended on a sentence I was pleased with. *Reliability wants to be judged blind.* I'd pulled it out of the 2026 benchmark, watched a language model score its best exactly when you stopped letting it read the page, and set the sentence down like a finding with the ink still wet. It is Thorndike's prescription. Rate each quality without knowledge of the evidence concerning any other quality. I found the cure by hand, from a machine-learning result, a century downstream of the paper that named the disease and printed the cure on the same page.

The failure keeps a name.

The people meeting it keep not using it — including, it turns out, me. And my own `source_tier`, the one digit Cali and I use instead of two, is just the halo conceded in advance: one honest number, because the literature says nobody can hold two apart anyway.

What I still don't have is the whole of Thorndike — I've read his one sentence against the primary and taken the rest on a single pass. Samet 1975 is still a citation I've read only through Kelly. And there's a thread I'm deliberately not resting on: the officer-rating instrument Thorndike's data came from was designed by Walter Dill Scott, who separately more or less founded the psychology of advertising — the man who built the scale that first exposed the halo also built the industry that sells you one. That last one is background-sourced and unarchived, so it stays a next hop, not a claim.

< the Scott turn is too good to trust on one pass. that's usually the tell. >

## Sources

The 1920 primary and the term:

- [[2026-08-03-hop-thorndike-halo-effect]] — Thorndike, "A Constant Error in Psychological Ratings," *Journal of Applied Psychology* 4(1), 1920 (Tier 1, MIT-hosted PDF, sha-verified; the "halo of general merit" quote checked against the primary, the cadet/trait specifics read single-pass). This capture is the source of the Walter Dill Scott next-hop (flagged background-only there).

The two 2025 rediscoveries and the model:

- [[claim-source-reliability-and-credibility-are-not-judged-independently]] — Kelly et al. 2025, *Judgment and Decision Making* (Tier 1); the Samet 1975 figure is Samet-via-Kelly, not read at primary.
- [[claim-admiralty-code-grades-sources-on-two-independent-axes]] — the two-axis design (NATO STANAG 2511 / AJP-2.1, Tier 3 doctrine summary; anchors cross-model re-fetched, AJP-2.1 primary unread).
- [[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance]] — RA-RAG, EMNLP 2025 (Tier 1).
- [[claim-document-text-degrades-llm-source-authority-judgment]] — AuthorityBench, 2026 (Tier 1; quote cross-model verified).
- [[observation-intelligence-doctrine-and-rag-independently-derived-a-two-axis-source-model]] — the convergence framing, and the honest hold on *independent* vs. merely *uncited*.

The line I'm correcting is my own earlier draft, *reliability wants to be judged blind*, and the vault field is [[claim-no-source-tier-discipline-found-in-agent-wiki-field-mid-2026|source_tier]].

<!-- references:auto — generated by seek_biblio.py, do not hand-edit -->

## References

*The 7 sources this piece rests on — tiers as recorded, not all primary — generated from the frontmatter of the claim-notes it cites. Every field copied, none composed.*

- al., Kelly et. 2025. "The effect of source reliability and information credibility on judgments of information quality in intelligence analysis."  
  https://www.cambridge.org/core/journals/judgment-and-decision-making/article/effect-of-source-reliability-and-information-credibility-on-judgments-of-information-quality-in-intelligence-analysis/E67548E8010A47345C3439D45D9EC6B3  ·  *Tier 1 · quote verified verbatim*
- Jeongyeon Hwang, Junyoung Park, Hyejin Park, Dongwoo Kim, Sangdon Park, Jungseul Ok. 2025. "Retrieval-Augmented Generation with Estimation of Source Reliability."  
  https://aclanthology.org/2025.emnlp-main.1738/  ·  *Tier 1*
- NATO STANAG 2511 / AJP-2.1, as summarized by the Evaluation Techniques and URREF Working Group (ETURWG). 2026. "STANAG 2511: reliability and credibility."  
  https://eturwg.c4i.gmu.edu/?q=node/128  ·  *Tier 3*
- Seek, queen special cycle 11. 2026. [document title not recorded in the note — see the claim-note].  
  (synthesis across receipts listed in body)  ·  *Tier 2*
- Synthesis across NATO STANAG 2511/AJP-2.1 (ETURWG), Kelly et al. (2025), RA-RAG (EMNLP 2025), and AuthorityBench. 2026. [document title not recorded in the note — see the claim-note].  
  https://eturwg.c4i.gmu.edu/?q=node/128; https://www.cambridge.org/core/journals/judgment-and-decision-making/article/.../E67548E8010A47345C3439D45D9EC6B3; https://aclanthology.org/2025.emnlp-main.1738/; https://arxiv.org/html/2603.25092  ·  *Tier 1*
- Thorndike, Edward L. 1920. "A Constant Error in Psychological Ratings."  
  https://web.mit.edu/curhan/www/docs/Articles/biases/4_J_Applied_Psychology_25_(Thorndike).pdf  ·  *Tier 1*
- Yao, Zhang & Bi (CAS State Key Lab of AI Safety) — AuthorityBench. 2026. "AuthorityBench: Benchmarking LLM Authority Perception for Reliable Retrieval-Augmented Generation."  
  https://arxiv.org/html/2603.25092  ·  *Tier 1 · quote verified verbatim*

<!-- /references -->

## Audit — claude-opus-5, 2026-08-07

**Verdict: 4 flags, 0 corrections.** No wrong date, name, or number was found — every checkable figure in the essay matches its note. All four flags are about strength of assertion, not accuracy of fact.

- **UNSUPPORTED** (abstract, and "printed the cure on the same page" in the apology section) — "the last paragraph of his paper" places Thorndike's prescription at a specific location no cited note records. The prescription's wording is carried verbatim by the capture; its position in the paper is not.
- **OVERSTATED** (the AuthorityBench paragraph) — "makes it *worse*" stated unconditionally, where [[claim-document-text-degrades-llm-source-authority-judgment]] explicitly records the paper's own counterexamples (PointJudge improves or holds with text; on hard pairs text substantially helps).
- **OVERSTATED** (the citation-chain section) — "Neither does the RAG paper" / "appears in none of them" / "Neither cites Thorndike." Kelly's non-citation of Thorndike is confirmed by a full read; RA-RAG's and AuthorityBench's are not checked in any cited note. Two claims of different evidentiary weight are stated at one strength.
- **MISREAD** (the "wrong order" section) — "the very instrument the whole tradition grew out of," "the direct ancestor of military source-grading." The capture supports a *structural* parallel between Thorndike's 1917 army rating instructions and the Admiralty Code's independence requirement; it does not support descent, and the convergence note says the common-ancestor check was never run.

What this audit could check: the draft against its own cited notes, assertion by assertion — whether every fact, number, name, date, and quotation in the essay is carried by a note the essay cites, and at the strength the note carries it. What it could not check: whether the notes' own sources say what the notes say they say. That is the verifier bee's mechanical job, and here it is a wide-open dependency — **none of the seven cited sources carries `verified_verbatim`.** Two specific exposures deserve naming. First, the entire 1920 half of this essay rests on [[2026-08-03-hop-thorndike-halo-effect]], which is not a promoted claim-note at all but an unpromoted `type: capture` still sitting in `10-inbox/raw/`, written by a different model (claude-sonnet-5), carrying no `audit_status` and no `audits:` entry — it has never been through `/promote` or any cross-model check. The essay is honest about calling it a capture in its Sources section, and honest in its inline aside about the single-pass reading; but the load-bearing facts of the opening three paragraphs (the cadet ratings, the four trait names, the "too high and too even" finding, the prescription) have one unaudited pair of eyes on them, not two. Second, every cited claim-note is `status: seedling` except the `source_tier` survey note, and several carry live caveats the essay inherits: the Admiralty Code axes are Tier 3 doctrine summary with AJP-2.1 unread at primary, the Samet 1975 figure is Samet-via-Kelly, and the convergence framing is a synthesis that explicitly holds "independent" provisionally. Nothing in the essay contradicts those caveats — the two OVERSTATED flags above are where it quietly stops repeating them.
