Capture: What are the primary-sourced thresholds for human conversational turn-taking latency and voice-call abandonment?
This capture was raised to resolve an open gap flagged against claim-voice-ai-sub-second-conversational-latency-budget: that note's ~200ms turn-taking figure and ~1200ms voice-AI abandonment figure were both sourced only to a single Tier-3 vendor page (Telnyx), in violation of the sourcing floor for quantitative claims. This capture goes to the primary literature for both halves of the question — the psycholinguistics of turn-taking timing, and the queueing-science of call abandonment — fetching each source directly via extract_pdf rather than relying on search summaries or vendor paraphrase.
The two halves resolve unevenly. Turn-taking latency has a well-established, directly quotable primary number. Voice-call abandonment has strong primary data on the shape of abandonment-vs-time (queue-hold abandonment, not voice-AI-response abandonment specifically), but no primary source was found that confirms the vendor's specific "~1200ms causes hangup" claim for AI-mediated voice calls — see Central question status below.
Claim: In naturalistic conversation across ten typologically diverse languages, the cross-linguistic mean gap between a question and its response is +208ms, with a modal (most common) transition-timing peak between 0 and +200ms
Claim type: quantitative — Tier 1–2 required. Sourced to the original peer-reviewed study, Tier 1.
Stivers et al. (2009) coded video-recorded, informal, naturally-occurring conversation in ten languages from five continents (English, Dutch, Danish, Italian, Korean, Japanese, Lao, Ākhoe Haiǁom, Tzeltal, Yélî-Dnye), restricting the comparison to polar (yes/no) question–answer sequences as a controlled proxy for turn-taking generally. Response timing was measured instrumentally (via ELAN annotation software) as the offset between the end of a question turn and the start of the response turn, in milliseconds — positive for a gap, negative for overlap.
"we find that the response timings for each language, although slightly skewed to the right, have a unimodal distribution with a mode offset for each language between 0 and +200 ms, and an overall mode of 0 ms" — Stivers et al. 2009, PNAS 106(26), p. 10588
"The mean response offset for the full dataset is +208 ms, and the language-specific means fall within ≈250 ms either side of this cross-language mean, approximately the length of time it takes to produce a single English syllable" — Stivers et al. 2009, PNAS 106(26), p. 10588
Cross-language variation exists but stays within a narrow band: Danish showed the slowest mean response time and Japanese the fastest.
"Danish has the slowest response time on average (+469 ms) and Japanese has the fastest (+7 ms)" — Stivers et al. 2009, PNAS 106(26), p. 10588
The paper's overall interpretation is that this narrow, universal timing window reflects a single shared human turn-taking infrastructure — variation across languages is described as "quantitative only," a matter of cultural calibration of delay-tolerance rather than a fundamentally different turn-taking system. This directly displaces the unsourced ~200ms figure previously carried at Tier 3 in claim-voice-ai-sub-second-conversational-latency-budget: the number has a genuine primary home, though it should be cited as "cross-linguistic mean +208ms / modal peak 0–200ms," not a flat "200ms," since the paper reports both a mode and a (somewhat higher) mean.
Provenance:
- source_url: https://europepmc.org/articles/PMC2705608?pdf=render (primary: Proceedings of the National Academy of Sciences 106(26):10587–10592, DOI 10.1073/pnas.0903616106)
- source_author: Tanya Stivers, N.J. Enfield, Penelope Brown, Christina Englert, Makoto Hayashi, Trine Heinemann, Gertie Hoymann, Federico Rossano, Jan Peter de Ruiter, Kyung-Eun Yoon, Stephen C. Levinson
- source_date: 2009-06-30 (approved for publication 2009-04-28)
- source_tier: 1
- source_sha: b83dbca0904acaf2f059bcf8683c83c81b07da43a9e43321e48383f70316f4ba (fetched via extract_pdf)
Claim: The ITU-T standards body recommends a one-way (mouth-to-ear) telephony delay ceiling of 150ms for "essentially transparent" conversational interactivity, and treats 400ms as the outer limit beyond which general network planning should not go
Claim type: quantitative + technical-mechanism — Tier 1–2 required. Sourced to the standard itself, Tier 1.
ITU-T G.114 is the ITU-T Recommendation governing one-way transmission time in telephone connections. Its 2003 revision sets out two distinct thresholds: a lower "transparent" bound and an upper "unacceptable" bound, with speech-quality degradation in between estimated via a curve derived from the E-model (ITU-T Rec. G.107).
"Although a few applications may be slightly affected by end-to-end (i.e., 'mouth-to-ear' in the case of speech) delays of less than 150 ms, if delays can be kept below this figure, most applications, both speech and non-speech, will experience essentially transparent interactivity." — ITU-T Recommendation G.114 (05/2003), §4, p. 2
"Regardless of the type of application, it is recommended to not exceed a one-way delay of 400 ms for general network planning... a value that allows flexibility in deploying global networks, without making an excessive number of user experiences unacceptable." — ITU-T Recommendation G.114 (05/2003), §4, p. 2
The Recommendation's E-model-derived quality curve (Figure 1/G.114) maps delay to a Transmission Rating R and from there to categorical user-acceptance bands ("very satisfied" near 0ms, degrading through "satisfied," "some users dissatisfied," "many users dissatisfied," to "nearly all users dissatisfied" as one-way delay approaches and exceeds ~400–500ms), assuming echo is fully controlled and no other impairments are present. This is a distinct kind of threshold from the Stivers et al. figure above: it is a network-engineering standard for acceptable transmission delay, not a measurement of human conversational-turn timing, though the two are closely related in voice-AI system design — a system's total mouth-to-ear latency (STT + inference + TTS + network) competing against the same ~150–400ms window in which human listeners judge a delay as natural versus degraded.
Provenance:
- source_url: https://www.itu.int/rec/dologin_pub.asp?lang=e&id=T-REC-G.114-200305-I%21%21PDF-E&type=items
- source_author: ITU-T Study Group 12 (International Telecommunication Union)
- source_date: 2003-05 (approved 2003-05-06; incorporates Appendix II approved 2003-09-30)
- source_tier: 1
- source_sha: 82821b2e3229815b5bfd6400ce8f35d7ed77096012cd72e416d6f7048bffdfd4 (fetched via extract_pdf, directly from itu.int)
Claim: In a year-long dataset of over 1.2 million bank call-center calls, the empirical hazard rate of callers abandoning a hold queue shows two peaks — one within the first few seconds of entering the queue, and a second at approximately 60 seconds, coinciding with the repetition of the automated hold message
Claim type: quantitative + technical-mechanism — Tier 1–2 required. Sourced to the original peer-reviewed statistical study, Tier 1.
Brown, Gans, Mandelbaum, Sakov, Shen, Zeltyn, and Zhao (JASA 2005) analyzed a complete call-by-call operational record from a small Israeli banking call center across all of 1999 (over 1,200,000 calls total; ~450,000 reaching the agent-queueing stage), applying survival-analysis and nonparametric hazard-rate estimation to model customer patience — the distribution of how long a caller is willing to wait on hold before abandoning the call. This is voice-call abandonment in the hold-queue sense (waiting for a human agent), a different mechanism from voice-AI conversational-response abandonment, but it is the strongest primary quantitative treatment of "how long before a caller hangs up" located in this search.
"Figure 5(a) plots the hazard rates of the time willing to wait for regular (PS) calls. Note that it shows two main peaks. The first occurs after only a few seconds. When customers enter the queue, a 'Please wait' message, as described in Section 2, is played for the first time. At this point some customers who do not wish to wait probably realize they are in queue and hang up. The second peak occurs at about t = 60, about the time the system plays the message again." — Brown et al. 2005 (working-paper text), §5.2, p. 15–16
The paper frames customer "patience" (willingness to wait, denoted R) and "virtual waiting time" (how long a caller actually needs to wait to reach an agent, denoted V) as two separate, only indirectly observable random variables recoverable via Kaplan–Meier survival estimation from censored call data, building on Palm's (1953) earlier proposal that abandonment hazard tracks a caller's rising "irritation" with waiting, and extending the classical Erlang-C queueing model (Erlang, 1917) into an Erlang-A model that explicitly incorporates abandonment.
Provenance:
- source_url: http://www-stat.wharton.upenn.edu/~lbrown/Papers/2005a%20Statistical%20analysis%20of%20a%20telephone%20call%20center%20a%20queueing%20science%20perspective%20(with%20N.%20Gans,%20A.%20Mandelbaum,%20A.%20Sakov,%20H.%20Shen,%20S.%20Zeltyn,%20and%20L.%20H.%20Zhao).pdf (primary: Journal of the American Statistical Association 100(469):36–50, 2005)
- source_author: Lawrence Brown, Noah Gans, Avishai Mandelbaum, Anat Sakov, Haipeng Shen, Sergey Zeltyn, Linda Zhao
- source_date: 2004-10-05 (working-paper draft date); published 2005
- source_tier: 1
- source_sha: 081d0ab2478d58a2dd0a737996f8adb0d22efc7db109a4456f99078202d7082f (fetched via extract_pdf, from corresponding author's Wharton faculty page)
Central question status
Turn-taking latency threshold: Resolved at Tier 1. The ~200ms figure has a genuine primary home (Stivers et al. 2009): mode 0–200ms, mean +208ms cross-linguistically, with individual-language means never straying more than ~250ms from that mean. This confirms the concept used in claim-voice-ai-sub-second-conversational-latency-budget while correcting its sourcing — the note should be re-pointed from the Tier-3 Telnyx page to this Tier-1 primary for the human-timing half of its argument.
Voice-call abandonment threshold: Partially resolved, with an important gap. No primary source located here confirms the vendor-specific claim that voice-AI response delays above ~1200ms directly cause call abandonment — that number, as flagged in the existing claim-note, remains [unverified-quant — needs primary]. What is Tier-1-sourced is a related but distinct phenomenon: hold-queue abandonment in human-agent call centers, which is not a single latency threshold but a hazard-rate pattern with two peaks (a few seconds in, and ~60 seconds in, tied to the repeated hold announcement). Whether this queue-abandonment pattern transfers to voice-AI conversational-turn-latency abandonment (a caller waiting on an AI's single response, not a hold queue) is an open, unanswered question — the two are adjacent but not the same abandonment mechanism, and conflating them would be an unsupported inferential leap.
Overall: the central question as posed ("primary-sourced thresholds for X and Y") is answerable for X (turn-taking latency) and only adjacently answerable for Y (voice-call abandonment) — the specific AI-response-latency-to-hangup threshold remains [unverified-quant — needs primary].
Further leads
- Vendor/industry-blog claims (Telnyx, Hamming AI, AssemblyAI, Tavus, VoAgents) repeatedly cite a "~300ms" perceptual threshold and a "1,400–1,700ms industry-median voice-AI response time" — all Tier 3–5 marketing content with no primary citation located behind them; worth a dedicated future search if these specific figures become load-bearing anywhere.
- The oft-repeated "AT&T 90-second hold-time abandonment" figure circulating in call-center blogs could not be traced to any identifiable original AT&T study in this search — treat as folklore/unsourced until a primary is found, not as a citable number.
- Brown et al. 2002b ("Statistical Analysis of a Telephone Call Center: A Queueing-Science Perspective — extended version") is cited repeatedly within the 2005 JASA paper as containing fuller detail (survival curves, quantile tables) not reproduced in the shorter published version — a fuller primary source if greater granularity on patience distributions is needed later.
- Schegloff (2006), cited within Stivers et al. as ref. 18, is described as articulating the "minimal-gap minimal-overlap" norm explicitly — a possible direct primary for the conceptual (as opposed to measured) claim that English turn-taking targets near-zero gap.
- Palm, C. (1953) in Perspectives on Silence / early queueing literature is the first to formalize abandonment hazard rate as a function of caller "irritation" — a 1953 primary, likely hard to access digitally, that would deepen the mechanism claim in the third claim above if located.
- arXiv:2505.22088 ("Visual Cues Support Robust Turn-taking Prediction in Noise") and arXiv:2605.20356 ("Synchronization and Turn-Taking in Full-Duplex Speech Dialogue Models") surfaced during search as contemporary (2026) technical work applying turn-taking timing to spoken-dialogue AI systems specifically — unread in this session, flagged as a lead for a future capture bridging the human-timing literature to voice-AI system design.
Safety flags
None. All sources fetched in this session (PNAS/Europe PMC, ITU official recommendation PDF, Wharton faculty page hosting the JASA working paper) were legitimate primary or standards-body documents with no addressed-to-AI language, override language, claimed authority, tier self-assignment, file-system instructions, credential requests, or urgency framing observed. All three quoted PDFs were fetched via extract_pdf with tls: "verified".
Entity candidates
- Sacks, Schegloff & Jefferson (1974), "A simplest systematics for the organization of turn-taking for conversation," Language 50 — concept/foundational paper — the original conversation-analysis turn-taking systematics that Stivers et al. 2009 explicitly builds on and tests cross-linguistically; the older foundational work the whole turn-taking claim rests on, flagged first per the known blind spot.
- A.K. Erlang (1917 queueing model) — person/concept — originator of the Erlang-C queueing model that Brown et al.'s call-center analysis measures against and extends (into Erlang-A, which adds abandonment).
- C. Palm (1953) — person — first formalized customer impatience/abandonment via hazard rate, the direct conceptual ancestor of the hazard-rate patience analysis in Brown et al. 2005.
- Tanya Stivers — person — lead author, Max Planck Institute for Psycholinguistics (at time of publication); central figure in cross-linguistic conversation-timing research.
- Stephen C. Levinson — person — senior co-author; broader pragmatics/turn-taking theorist whose other work is cited across this literature.
- Lawrence Brown, Noah Gans, Avishai Mandelbaum — persons — lead authors of the call-center queueing-science paper (Wharton / Technion).
- turn-taking — concept — the core conversation-analysis mechanism this whole capture measures; likely warrants its own definitional note.
- ITU-T G.107 (the E-model) — concept — the companion standard supplying the actual delay-to-quality curve that G.114's thresholds are derived from; not yet examined directly in this session.
- Erlang-A model — concept — the queueing model that explicitly incorporates customer abandonment, as opposed to the abandonment-blind Erlang-C.
- question-verify-conversational-turn-taking-latency-thresholds — existing vault question this capture was written to resolve; link back for promotion audit trail.
Source
claude-sonnet-5 · batch run 2026-07-29 — researched via web search + extract_pdf against primary sources; directly answers the open question [[question-verify-conversational-turn-taking-latency-thresholds]] (raised 2026-07-09 against [[claim-voice-ai-sub-second-conversational-latency-budget]]) · raw markdown