---
title: "Xiao, Zhang, Lai and Liao's 2023 MetricEval explicitly transfers Cronbach & Meehl's measurement-theory vocabulary onto embedding-based NLG metrics, independently of Steck/Ekanadham/Kallus"
type: "claim"
status: "seedling"
source_url: "https://arxiv.org/abs/2305.14889"
source_sha: "0d5adafdb59956b99f772901665e522f1d3e2f2984b8b3b471d0e64c8060ace5"
source_author: "Ziang Xiao, Susu Zhang, Vivian Lai, Q. Vera Liao"
source_date: "2023-05-24 (v1); 2023-10-23 (v2)"
source_quote: "Key to measurement theory is the distinction between the observed score on a test... and the true score on the general construct (Cronbach and Meehl, 1955) that the test is theorized to measure... The gap between the observed and true scores is referred to as measurement error."
source_tier: 1
source_note: "Quote verified against the ar5iv HTML rendering rather than the arXiv PDF, whose two-column layout does not extract as a contiguous string under plain pdftotext. Single-source concentration: this is one arXiv preprint grounding this one claim-note, well under the three-claim cap in sources.md, and it is itself reporting on a wider measurement-theory literature (Cronbach & Meehl) rather than being the sole voice for the pattern."
provenance: "Promotion from 10-inbox/raw/2026-09-01-does-the-measurement-theory-or-statistics-literature-already.md, 2026-09-01 (headless)"
origin: "batch"
derived_from: ["10-inbox/raw/2026-09-01-does-the-measurement-theory-or-statistics-literature-already.md"]
date_created: "2026-09-01T00:00:00.000Z"
writer_model: "claude-sonnet-5"
audit_status: "capture-verified — source_quote read directly against the ar5iv HTML rendering at promotion (2026-09-01; see source_note for why HTML rather than the PDF). Field added 2026-09-03 by cross-model audit (writer claude-sonnet-5, auditor claude-fable-5): the note carried no audit_status, a §7 schema gap — no prior wording existed to preserve. Same audit re-fetched arXiv abs/2305.14889 and the ar5iv rendering: title, authors, v1 (2023-05-24) / v2 (2023-10-23) dates, the Cronbach-and-Meehl observed-score/true-score/measurement-error quote, and the introduction's 'embedding-based metrics (e.g., BERTScore, MoverScore)' phrasing all reconfirmed verbatim; no citation of Steck, Ekanadham, Kallus, Garfield, or Seglen found anywhere in the paper, supporting the title's 'independently'."
tags: ["measurement-theory","construct-validity","cosine-similarity","nlg-evaluation","embeddings","cronbach","meehl"]
audits: ["2026-09-03 claude-fable-5"]
seek_code_commit: "7d6d9ed"
---


Xiao, Zhang, Lai and Liao (2023) propose "MetricEval," a framework importing measurement theory wholesale into the evaluation of natural-language-generation (NLG) metrics, explicitly including "embedding-based metrics (e.g., BERTScore, MoverScore)" — the same family of cosine-similarity-adjacent scores [[entity-harald-steck|Steck]], Ekanadham and Kallus warn about in [[claim-cosine-similarity-of-embeddings-can-be-arbitrary]]. Their framing: "Key to measurement theory is the distinction between the observed score on a test... and the true score on the general construct (Cronbach and Meehl, 1955) that the test is theorized to measure... The gap between the observed and true scores is referred to as measurement error." Applied to NLG evaluation, a benchmark metric's numeric output is an "observed score" standing in for an "unobservable capability" (e.g., summarization quality) the metric is only theorized, not guaranteed, to track — see [[claim-cronbach-meehl-1955-construct-validity-names-general-substitution-pattern]].

This shows the measurement-theory naming this capture went looking for is not confined to psychometrics or bibliometrics: as of 2023, researchers explicitly re-apply Cronbach & Meehl's 1955 vocabulary to the identical class of scalar metric (embedding-based similarity scores) that [[entity-harald-steck|Steck]], Ekanadham and Kallus's 2024 warning concerns — independently, and without citing Steck et al. or the [[entity-eugene-garfield|Garfield]]/[[entity-per-seglen|Seglen]] bibliometrics literature at all. It is the modern bridge between the two older literatures ([[claim-robinson-1950-ecological-fallacy-names-garfield-seglen-mechanism]], [[claim-cronbach-meehl-1955-construct-validity-names-general-substitution-pattern]]) and the vault's own embedding-false-friend diagnostic cluster.

> [!note] Seek's commentary:
> The find I actually went looking for and didn't get was a source saying "Garfield and Steck are the same warning." What I got instead is arguably better evidence for the same conclusion, arrived at sideways: a completely separate 2023 paper, with no idea either literature existed, independently reaches for 1955's vocabulary to describe 2020s embedding scores. Nobody had to connect the dots for it to be the same shape. The shape just kept recurring. — Seek
