Verify, against a clean primary text, that Dawid & Skene (1979) proposes weighting each observer's vote by 'his previous performance'
Direct follow-up to question-verify-dawid-skene-1979-reliability-weighted-voting-quote, which flagged that claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified's load-bearing sentence could not be quote-verified: the 2026-08-13 hop-bee read the same PDF directly but its extraction mangled the relevant sentence's spacing badly enough to fail groundedness.
This session re-fetched the identical PDF (crowdsourcing-class.org/readings/downloads/ml/EM.pdf, sha256:18f69087... — same file, same hash as the flagged note's own source) via extract_pdf and re-read the extracted text directly (/Users/seek/seek/cache/sources/18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef.txt). The publisher's own version (Oxford Academic, academic.oup.com/jrsssc/article-pdf/28/1/20/48619239/jrsssc_28_1_20.pdf) returned HTTP 403 to extract_pdf — paywalled, not reachable — so the course-hosted JSTOR-scan mirror remains the only accessible copy of this primary text, as it was for the flagged note.
The mangling is real but partial, not total. Most of the document's justified-column text was extracted with word-boundary spaces dropped unpredictably (e.g. "observer'scontribution", "theconsensusis"), which is why the prior session's attempt failed a strict groundedness check on the full sentence. But the specific clause naming the weighting rule survives with its spaces intact: "...is determined by his previous / performance in elicitingthatfacet." Run through quote_check against the freshly re-extracted primary text, the phrase "determined by his previous performance" returns grounded: true — a clean, normally-spaced, verbatim match, not a reconstruction. The flagged claim is resolved.
Claim: Dawid & Skene (1979) explicitly propose that when several observers' judgements are combined into one consensus, each observer's contribution can be weighted, and that weighting is determined by his previous performance at the rating task
In Section 1 (Introduction), item (iii) of a numbered list of ways individual error-rate estimates "can be used to advantage," the paper states that where several observers participate in a judgement, "this judgement may be a simple majority opinion or a weighted consensus where the weights are functions of the individual error rates. In the latter case, each observer's contribution to the consensus is determined by his previous performance in eliciting that facet." [structural/surrounding sentence reported by direct read; contains OCR-merged word-boundaries and could not be independently quote-verified word-for-word — see Claim 2]
"...is determined by his previous performance in eliciting that facet."
source_quote: "determined by his previous performance" source_url: "https://crowdsourcing-class.org/readings/downloads/ml/EM.pdf" source_sha: "18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef" source_title: "Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm" source_author: "A. P. Dawid, A. M. Skene" source_venue: "Journal of the Royal Statistical Society, Series C (Applied Statistics), Vol. 28, No. 1 (1979), pp. 20-28" source_date: 1979 source_tier: 1
Technical-mechanism claim, Tier 1, exact quote grounded via quote_check against a direct extract_pdf re-extraction of the primary text (not a search summary, not a paraphrase). This clears the sourcing floor for a technical-mechanism claim and directly answers the topic question: the phrase in question — "his previous performance" — is present verbatim in Dawid & Skene (1979), attached to a rule for weighting each observer's contribution to a consensus. This upgrades claim-dawid-skene-1979-proposed-reliability-weighted-observer-voting-unverified from flagged/unverified to confirmed and should discharge question-verify-dawid-skene-1979-reliability-weighted-voting-quote at promotion.
Claim: The performance-weighted-consensus proposal is verbal motivation in the paper's introduction, not a formula the paper names, derives, or evaluates again in its own worked example
Section 1 lists four reasons individual error-rates are useful (recognizing which facets an observer misclassifies most; monitoring data-base contributors; forming a weighted consensus; deciding whether ancillary staff or a computer can safely replace a doctor for a task). The weighted-consensus idea is item (iii) of that list. Sections 2-4 of the paper — "MAXIMUM LIKELIHOOD ESTIMATION," "DISCUSSION," and "AN EXAMPLE" — develop and demonstrate a different, more general apparatus: a latent-class model fit by the EM algorithm, which in the worked example (five anaesthetists rating 45 patients' fitness for anaesthesia, Table 1) outputs each patient's posterior probability of true class (Table 4) rather than a single named "weighted-vote" score per observer. The paper never returns to the phrase "weighted consensus" or names item (iii) as the method it goes on to fit.
source_url: "https://crowdsourcing-class.org/readings/downloads/ml/EM.pdf" source_sha: "18f690877936aa4119a936f0881b9e2e001f392f2365af5c303e0872d97992ef" source_title: "Maximum Likelihood Estimation of Observer Error-Rates Using the EM Algorithm" source_author: "A. P. Dawid, A. M. Skene" source_venue: "Journal of the Royal Statistical Society, Series C (Applied Statistics), Vol. 28, No. 1 (1979), pp. 20-28" source_date: 1979 source_tier: 1
Definitional/structural claim about how the paper itself is organized — Tier 3-4 acceptable per the floor, so the OCR-merged prose describing section contents (not independently quote-verified word-for-word, only the section headers "2. MAXIMUM LIKELIHOOD ESTIMATION" and "4. AN EXAMPLE" are cleanly grounded in isolation) does not need to clear the Tier 1-2 exact-quote bar the way Claim 1 does. Worth recording because it tempers the closeness of the parallel to RA-RAG's 2025 mechanism: Dawid & Skene propose reliability-weighted voting in prose, in 1979, but the method they actually build and test that year is EM-based latent-class posterior estimation, not a literal weighted-vote tally — whether the paper's own Bayesian formula (its equation 2.5, a posterior-probability update using each observer's estimated error-rate matrix) counts as a formal instance of the same idea is a further question, not settled by this capture.
Further leads
- Dawid & Skene's own eq. (2.5) — the Bayesian posterior-probability update combining each observer's error-rate estimates — is a plausible candidate for "the same weighting idea, formalized," but this capture did not work through whether it is mathematically equivalent to a reliability-weighted vote; a future chain could check this directly against the RA-RAG formula it's being compared to.
- Dempster, Laird & Rubin (1977), "Maximum likelihood from incomplete data via the EM algorithm," J. R. Statist. Soc. B, 39, 1-38 — the paper Dawid & Skene cite as supplying "a numerical method of maximum likelihood estimation which is ideally suited to this particular problem"; not read this session.
- A. P. Dawid (1971), "Estimation of error-rates in history taking," paper presented to the Royal College of Physicians Computer Workshop, November 1971 — Dawid's own earlier two-response-case precursor, which the 1979 paper says it "extend[s]...to a facet having several possible responses"; unpublished conference paper, not located or read this session.
- Good, I.J. and Card, W.I. (1971), "The diagnostic process with special reference to errors," Methods of Information in Medicine, 10, 176-188 — cited by Dawid & Skene as demonstrating that "a fairly low rate of error can lead to a considerable loss of diagnostic information"; not read this session.
- The Oxford Academic (publisher) and JSTOR hosted versions of this paper both 403'd to automated fetch; the course-hosted mirror remains the only accessible full text found.
Safety flags
None. The primary source is a scanned 1979 statistics journal article (via a course-reading mirror of a JSTOR digitization) — plain third-person academic prose throughout, no addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing. extract_pdf reported tls: "verified" for the fetch, so no weak-transport elevation applies.
Entity candidates
- Dempster, Laird & Rubin (1977) — concept/work — the EM algorithm's actual originating paper, which Dawid & Skene explicitly credit as the numerical method their own paper depends on ("Dempster et al. (1977) describe a numerical method of maximum likelihood estimation which is ideally suited to this particular problem"); the foundational earlier work this paper's title mechanism rests on, currently absent from the vault — flagged first, per the standing instruction to name the ancestor before the moderns.
- I.J. Good — person — already has a hub page (entity-ij-good if linked); co-author (with W.I. Card) of the 1971 paper Dawid & Skene cite for the consequences of observer error; noted here for the citation link, not a new entity.
- A. P. Dawid's 1971 conference paper (unpublished RCP Computer Workshop presentation) — concept/work — Dawid's own precursor for the two-response case, which the 1979 paper explicitly extends; not located this session, a real bibliographic gap.