---
id: "20260814-0215-verify-against-a-clean"
title: "Verify, against a clean primary text, that Dawid & Skene (1979) proposes weighting each observer's vote by 'his previous performance'"
type: "capture"
status: "promoted"
origin: "batch"
promoted_to: ["30-notes/claim-dawid-skene-1979-proposes-but-never-implements-weighted-consensus.md","30-notes/claim-dawid-skene-1979-credits-dempster-laird-rubin-1977-for-em-method.md","40-entities/entity-dempster-laird-rubin-em-algorithm.md (new entity hub)","30-notes/claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified.md (updated in place — corroboration line appended to audit_status, status and flag left untouched)",{"50-questions/question-verify-dawid-skene-1979-reliability-weighted-voting-quote.md (updated in place — dated progress log appended, left status":"open)"}]
not_promoted: ["Claim 1 (quote_check-grounded 'determined by his previous performance'): not written as a standalone new claim-note. It duplicates 30-notes/claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified.md, whose discharge is already escalated to Cali in 90-feedback/2026-08-14-from-auditor-dawid-skene-previous-performance-quote-recoverable.md because 70-drafts/filed-under-medical-example/draft.md (unapproved) rests a paragraph and a voice bracket on that flag staying open. A promotion pass has no more standing than that audit did to apply the discharge; forking a second 'confirmed' note would have left the vault holding two contradictory statuses on the same fact. Folded as a corroborating line into the existing note's audit_status and the question's progress log instead.","I.J. Good & W.I. Card (1971), cited by D&S as evidence that low observer error can still cause considerable diagnostic information loss — flagged in the capture's own Entity candidates section as 'noted here for the citation link, not a new entity'; the 1971 paper itself was not read this session, no claim-note backs the fact, and entity-ij-good.md already covers what the vault currently knows about Good, so it was left untouched rather than given a thin unbacked update line.","A. P. Dawid's 1971 unpublished RCP Computer Workshop paper (the two-response-case precursor D&S say they extend): real bibliographic gap, but unlocated and unread. UNSURE per the entity-promotion test — no page, no stub; left as a further lead.","Dempster, Laird & Rubin (1977)'s own paper content: not read this session, only the sentence in which D&S cite it. The claim promoted is 'D&S credit this paper for their method,' not an independent claim about DLR's own content — the entity hub records DLR's significance from what is already established about it, not from a primary read this session.","The Oxford Academic/JSTOR 403-on-fetch access note: a tooling/access fact, already recorded in the 2026-08-13 capture's further leads and the sibling flagged claim-note; not a new finding worth its own record."]
writer_model: "claude-sonnet-5"
date_created: "2026-08-14T00:00:00.000Z"
provenance: "this batch run, 2026-08-14"
derived_from: []
tags: ["dawid-skene","crowdsourcing","voting-theory","machine-learning-theory","medical-statistics","EM-algorithm","source-verification","RAG"]
source_url: "https://crowdsourcing-class.org/readings/downloads/ml/EM.pdf"
source_sha: "18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef"
source_title: "Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm"
source_author: "A. P. Dawid, A. M. Skene"
source_date: 1979
source_venue: "Journal of the Royal Statistical Society, Series C (Applied Statistics), Vol. 28, No. 1 (1979), pp. 20-28, Wiley for the Royal Statistical Society"
source_tier: 1
seek_code_commit: "17d9798"
---


Direct follow-up to [[question-verify-dawid-skene-1979-reliability-weighted-voting-quote]], which flagged that [[claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified]]'s load-bearing sentence could not be quote-verified: the 2026-08-13 hop-bee read the same PDF directly but its extraction mangled the relevant sentence's spacing badly enough to fail groundedness.

This session re-fetched the identical PDF (`crowdsourcing-class.org/readings/downloads/ml/EM.pdf`, `sha256:18f69087...` — same file, same hash as the flagged note's own source) via `extract_pdf` and re-read the extracted text directly (`/Users/seek/seek/cache/sources/18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef.txt`). The publisher's own version (Oxford Academic, `academic.oup.com/jrsssc/article-pdf/28/1/20/48619239/jrsssc_28_1_20.pdf`) returned HTTP 403 to `extract_pdf` — paywalled, not reachable — so the course-hosted JSTOR-scan mirror remains the only accessible copy of this primary text, as it was for the flagged note.

**The mangling is real but partial, not total.** Most of the document's justified-column text was extracted with word-boundary spaces dropped unpredictably (e.g. "observer'scontribution", "theconsensusis"), which is why the prior session's attempt failed a strict groundedness check on the full sentence. But the specific clause naming the weighting rule survives with its spaces intact: "...is determined  by his previous / performance    in elicitingthatfacet." Run through `quote_check` against the freshly re-extracted primary text, the phrase **"determined by his previous performance"** returns `grounded: true` — a clean, normally-spaced, verbatim match, not a reconstruction. The flagged claim is resolved.

## Claim: Dawid & Skene (1979) explicitly propose that when several observers' judgements are combined into one consensus, each observer's contribution can be weighted, and that weighting is determined by his previous performance at the rating task

In Section 1 (Introduction), item (iii) of a numbered list of ways individual error-rate estimates "can be used to advantage," the paper states that where several observers participate in a judgement, "this judgement may be a simple majority opinion or a weighted consensus where the weights are functions of the individual error rates. In the latter case, each observer's contribution to the consensus is determined by his previous performance in eliciting that facet." [structural/surrounding sentence reported by direct read; contains OCR-merged word-boundaries and could not be independently quote-verified word-for-word — see Claim 2]

> "...is determined by his previous performance in eliciting that facet."

source_quote: "determined by his previous performance"
source_url: "https://crowdsourcing-class.org/readings/downloads/ml/EM.pdf"
source_sha: "18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef"
source_title: "Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm"
source_author: "A. P. Dawid, A. M. Skene"
source_venue: "Journal of the Royal Statistical Society, Series C (Applied Statistics), Vol. 28, No. 1 (1979), pp. 20-28"
source_date: 1979
source_tier: 1

Technical-mechanism claim, Tier 1, exact quote grounded via `quote_check` against a direct `extract_pdf` re-extraction of the primary text (not a search summary, not a paraphrase). This clears the sourcing floor for a technical-mechanism claim and directly answers the topic question: the phrase in question — "his previous performance" — is present verbatim in Dawid & Skene (1979), attached to a rule for weighting each observer's contribution to a consensus. **This upgrades [[claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified]] from flagged/unverified to confirmed** and should discharge [[question-verify-dawid-skene-1979-reliability-weighted-voting-quote]] at promotion.

## Claim: The performance-weighted-consensus proposal is verbal motivation in the paper's introduction, not a formula the paper names, derives, or evaluates again in its own worked example

Section 1 lists four reasons individual error-rates are useful (recognizing which facets an observer misclassifies most; monitoring data-base contributors; forming a weighted consensus; deciding whether ancillary staff or a computer can safely replace a doctor for a task). The weighted-consensus idea is item (iii) of that list. Sections 2-4 of the paper — "MAXIMUM LIKELIHOOD ESTIMATION," "DISCUSSION," and "AN EXAMPLE" — develop and demonstrate a different, more general apparatus: a latent-class model fit by the EM algorithm, which in the worked example (five anaesthetists rating 45 patients' fitness for anaesthesia, Table 1) outputs each patient's *posterior probability* of true class (Table 4) rather than a single named "weighted-vote" score per observer. The paper never returns to the phrase "weighted consensus" or names item (iii) as the method it goes on to fit.

source_url: "https://crowdsourcing-class.org/readings/downloads/ml/EM.pdf"
source_sha: "18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef"
source_title: "Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm"
source_author: "A. P. Dawid, A. M. Skene"
source_venue: "Journal of the Royal Statistical Society, Series C (Applied Statistics), Vol. 28, No. 1 (1979), pp. 20-28"
source_date: 1979
source_tier: 1

Definitional/structural claim about how the paper itself is organized — Tier 3-4 acceptable per the floor, so the OCR-merged prose describing section contents (not independently quote-verified word-for-word, only the section headers "2. MAXIMUM LIKELIHOOD ESTIMATION" and "4. AN EXAMPLE" are cleanly grounded in isolation) does not need to clear the Tier 1-2 exact-quote bar the way Claim 1 does. Worth recording because it tempers the closeness of the parallel to [[claim-reliability-aware-rag-estimates-source-reliability-separately-from-relevance|RA-RAG's 2025 mechanism]]: Dawid & Skene *propose* reliability-weighted voting in prose, in 1979, but the method they actually build and test that year is EM-based latent-class posterior estimation, not a literal weighted-vote tally — whether the paper's own Bayesian formula (its equation 2.5, a posterior-probability update using each observer's estimated error-rate matrix) counts as a formal instance of the same idea is a further question, not settled by this capture.

> [!note] Seek's commentary:
> The prior session's flag was the right call — writing "determined by his previous performance" as a clean quote off a garbled extraction would have been exactly the kind of tidied-up fabrication risk the sourcing floor exists to prevent. What actually closed the gap wasn't a better PDF; it was rereading the *same* extraction and noticing that this one clause, unlike most of the paragraph around it, happened to keep its spaces. The lesson isn't "OCR problems resolve themselves" — most of this document is still unquotable — it's that a groundedness failure on a whole sentence doesn't mean every word-span inside it fails too, and it's worth testing sub-spans before concluding a document is a dead end.
> — Seek

## Further leads

- Dawid & Skene's own eq. (2.5) — the Bayesian posterior-probability update combining each observer's error-rate estimates — is a plausible candidate for "the same weighting idea, formalized," but this capture did not work through whether it is mathematically equivalent to a reliability-weighted vote; a future chain could check this directly against the RA-RAG formula it's being compared to.
- Dempster, Laird & Rubin (1977), "Maximum likelihood from incomplete data via the EM algorithm," *J. R. Statist. Soc. B*, 39, 1-38 — the paper Dawid & Skene cite as supplying "a numerical method of maximum likelihood estimation which is ideally suited to this particular problem"; not read this session.
- A. P. Dawid (1971), "Estimation of error-rates in history taking," paper presented to the Royal College of Physicians Computer Workshop, November 1971 — Dawid's own earlier two-response-case precursor, which the 1979 paper says it "extend[s]...to a facet having several possible responses"; unpublished conference paper, not located or read this session.
- Good, I.J. and Card, W.I. (1971), "The diagnostic process with special reference to errors," *Methods of Information in Medicine*, 10, 176-188 — cited by Dawid & Skene as demonstrating that "a fairly low rate of error can lead to a considerable loss of diagnostic information"; not read this session.
- The Oxford Academic (publisher) and JSTOR hosted versions of this paper both 403'd to automated fetch; the course-hosted mirror remains the only accessible full text found.

## Safety flags

None. The primary source is a scanned 1979 statistics journal article (via a course-reading mirror of a JSTOR digitization) — plain third-person academic prose throughout, no addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing. `extract_pdf` reported `tls: "verified"` for the fetch, so no weak-transport elevation applies.

## Entity candidates

- Dempster, Laird & Rubin (1977) — concept/work — the EM algorithm's actual originating paper, which Dawid & Skene explicitly credit as the numerical method their own paper depends on ("Dempster et al. (1977) describe a numerical method of maximum likelihood estimation which is ideally suited to this particular problem"); the foundational earlier work this paper's title mechanism rests on, currently absent from the vault — flagged first, per the standing instruction to name the ancestor before the moderns.
- I.J. Good — person — already has a hub page ([[entity-ij-good.md|entity-ij-good]] if linked); co-author (with W.I. Card) of the 1971 paper Dawid & Skene cite for the consequences of observer error; noted here for the citation link, not a new entity.
- A. P. Dawid's 1971 conference paper (unpublished RCP Computer Workshop presentation) — concept/work — Dawid's own precursor for the two-response case, which the 1979 paper explicitly extends; not located this session, a real bibliographic gap.
