Dawid & Skene (1979) propose performance-weighted consensus voting in their introduction but never name, derive, or test it as a method in the rest of the paper
Dawid & Skene (1979)'s Section 1 lists four practical uses for individual observer error-rate estimates. Item (iii) is a weighted consensus: "a weighted consensus where the weights are functions of the individual error rates," each observer's contribution "determined by his previous performance." But the paper's own body — Section 2 ("Maximum Likelihood Estimation"), Section 3 ("Discussion"), and Section 4 ("An Example," the five-anaesthetist worked case in Table 1) — develops and demonstrates a different apparatus: a latent-class model fit by the Dempster-Laird-Rubin EM algorithm, which outputs each patient's posterior probability of true class rather than a single weighted-vote score per observer. The paper never returns to the phrase "weighted consensus," and never names item (iii) as the method it goes on to fit.
This matters to the vault's citation-lineage thread because it tempers the closeness of the parallel to RA-RAG's 2025 mechanism: Dawid & Skene propose reliability-weighted voting in prose, in 1979, but the method they actually build and test that year is EM-based posterior estimation, not a literal weighted-vote tally. Whether the paper's own Bayesian update (its equation 2.5, combining each observer's estimated error-rate matrix into a posterior) is mathematically equivalent to a reliability-weighted vote is a separate, unresolved question — this note establishes only that the paper itself never makes that equivalence explicit.
Source
“2. MAXIMUM LIKELIHOOD ESTIMATION / 4. AN EXAMPLE”
claude-sonnet-5 · Promotion from 10-inbox/raw/2026-08-14-verify-against-a-clean-primary-text-that-dawid.md, 2026-08-14 · raw markdown