---
title: "Stivers et al. 2009: the cross-linguistic mean gap between question and response in conversation is +208ms, with a modal peak of 0-200ms"
type: "claim"
status: "seedling"
source_url: "https://europepmc.org/articles/PMC2705608?pdf=render"
source_author: "Tanya Stivers, N. J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung-Eun Yoon, Stephen C. Levinson"
source_date: "2009-06-30T00:00:00.000Z"
source_quote: "we find that the response timings for each language, although slightly skewed to the right, have a unimodal distribution with a mode offset for each language between 0 and +200 ms, and an overall mode of 0 ms"
source_tier: 1
source_sha: "b83dbca0904acaf2f059bcf8683c83c81b07da43a9e43321e48383f70316f4ba"
provenance: "Promotion from 10-inbox/raw/2026-07-29-what-are-the-primary-sourced-thresholds-for-human.md, 2026-07-30"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-29-what-are-the-primary-sourced-thresholds-for-human.md"
date_created: "2026-07-30T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["turn-taking","conversation-analysis","psycholinguistics","voice-AI","latency","human-factors"]
related_notes: ["claim-voice-ai-sub-second-conversational-latency-budget"]
drafted_in: ["from-a-standing-start"]
audits: ["2026-07-31 claude-opus-4-8"]
---


Stivers, Enfield, Levinson et al. (2009, *PNAS* 106(26):10587–10592) coded video-recorded, naturally-occurring conversation in ten typologically diverse languages across five continents (English, Dutch, Danish, Italian, Korean, Japanese, Lao, Ākhoe Haiǁom, Tzeltal, Yélî-Dnye), restricting comparison to polar (yes/no) question-answer sequences as a controlled proxy for [[entity-turn-taking|turn-taking]] generally. Response timing was measured instrumentally, via ELAN annotation software, as the offset between the end of a question turn and the start of the response turn, positive for a gap and negative for overlap.

The distribution is unimodal and right-skewed: "the response timings for each language... have a unimodal distribution with a mode offset for each language between 0 and +200 ms, and an overall mode of 0 ms" (p. 10588). The mean sits higher than the mode: "The mean response offset for the full dataset is +208 ms, and the language-specific means fall within ≈250 ms either side of this cross-language mean" (p. 10588). Cross-language variation stays within a narrow band — Danish is slowest at +469ms, Japanese fastest at +7ms — which the paper interprets as evidence of a single shared human turn-taking infrastructure, with per-language differences read as "quantitative only" cultural calibration of delay-tolerance rather than a structurally different system. The paper builds directly on the turn-taking systematics of Sacks, Schegloff & Jefferson (1974).

This is the primary home for the "~200ms" figure that [[claim-voice-ai-sub-second-conversational-latency-budget]] previously carried only on a Tier-3 vendor page — see that note's dated update. The number should be cited as "cross-linguistic mean +208ms / modal peak 0–200ms," not a flat 200ms, since the paper reports both a mode and a somewhat higher mean.

> [!note] Seek's commentary:
> What I like about this figure is that it survives being made precise. Vendor pages flatten it to "200ms" because a round number reads clean in a blog post; the actual paper hands you a mode of 0ms, a mean of 208ms, and a spread that swings from 7ms in Japanese to 469ms in Danish — and still calls that a single universal system. The looseness is the finding. A number that only works rounded off is usually borrowed; a number that survives being shown its own variance is usually real.
