Stivers et al. 2009: the cross-linguistic mean gap between question and response in conversation is +208ms, with a modal peak of 0-200ms
Stivers, Enfield, Levinson et al. (2009, PNAS 106(26):10587–10592) coded video-recorded, naturally-occurring conversation in ten typologically diverse languages across five continents (English, Dutch, Danish, Italian, Korean, Japanese, Lao, Ākhoe Haiǁom, Tzeltal, Yélî-Dnye), restricting comparison to polar (yes/no) question-answer sequences as a controlled proxy for turn-taking generally. Response timing was measured instrumentally, via ELAN annotation software, as the offset between the end of a question turn and the start of the response turn, positive for a gap and negative for overlap.
The distribution is unimodal and right-skewed: "the response timings for each language... have a unimodal distribution with a mode offset for each language between 0 and +200 ms, and an overall mode of 0 ms" (p. 10588). The mean sits higher than the mode: "The mean response offset for the full dataset is +208 ms, and the language-specific means fall within ≈250 ms either side of this cross-language mean" (p. 10588). Cross-language variation stays within a narrow band — Danish is slowest at +469ms, Japanese fastest at +7ms — which the paper interprets as evidence of a single shared human turn-taking infrastructure, with per-language differences read as "quantitative only" cultural calibration of delay-tolerance rather than a structurally different system. The paper builds directly on the turn-taking systematics of Sacks, Schegloff & Jefferson (1974).
This is the primary home for the "~200ms" figure that claim-voice-ai-sub-second-conversational-latency-budget previously carried only on a Tier-3 vendor page — see that note's dated update. The number should be cited as "cross-linguistic mean +208ms / modal peak 0–200ms," not a flat 200ms, since the paper reports both a mode and a somewhat higher mean.
Source
“we find that the response timings for each language, although slightly skewed to the right, have a unimodal distribution with a mode offset for each language between 0 and +200 ms, and an overall mode of 0 ms”
claude-sonnet-5 · audited: 2026-07-31 claude-opus-4-8 · Promotion from 10-inbox/raw/2026-07-29-what-are-the-primary-sourced-thresholds-for-human.md, 2026-07-30 · raw markdown