What are the primary-sourced thresholds for human conversational turn-taking latency and voice-call abandonment?
claim-voice-ai-sub-second-conversational-latency-budget states, on a single Tier-3 vendor page, that humans respond within ~200ms in conversation, that users hang up above ~1200ms, and that multi-vendor voice stacks add ≥250ms in network hops for an 800ms–1.5s total. Quantitative claims require Tier 1–2 sourcing; the note carries an [unverified-quant] flag and stays at seedling until the figures have primary homes.
Why it matters
The whole "voice AI must fit a sub-second budget" argument turns on these numbers. The ~200ms turn-taking figure in particular is treated as a perceptual constant, and it drives the case for colocated inference (claim-telnyx-colocates-gpus-at-telephony-pops-for-voice-latency). A concept resting on borrowed vendor digits is fragile; the concept is more defensible if the timing floor is anchored to the research literature rather than a marketing page.
What would answer it
- Turn-taking latency (~200ms): conversation-analysis / psycholinguistics literature on the timing of turn transitions (e.g. Stivers et al. 2009, PNAS, on universals in turn-taking timing, and related gap-duration studies). Find the actual measured inter-turn gap distributions.
- Call-abandonment threshold (~1200ms): telephony / contact-center human-factors studies on response-delay tolerance, rather than a vendor assertion.
- Multi-vendor hop cost (~250ms): an independent measurement or teardown of a chained STT→LLM→TTS stack, not the competitor-displacing vendor's own figure.
Candidate next moves
- Pull the turn-taking timing primary literature first — it is the most citable and least contested piece.
- Keep the vendor's <1s self-report clearly labelled as a self-report even if the human-side thresholds check out.
Progress log
2026-07-30 — partially answered. The turn-taking half is resolved: claim-stivers-2009-cross-linguistic-turn-taking-gap-208ms-mean gives the ~200ms figure a genuine Tier-1 primary home (mode 0-200ms, cross-linguistic mean +208ms, Stivers et al. 2009, PNAS), and claim-itu-t-g114-150ms-transparent-400ms-outer-limit-thresholds adds the adjacent standards-body delay ceiling (ITU-T G.114: 150ms transparent / 400ms outer limit). The abandonment half remains open: claim-brown-2005-call-center-abandonment-hazard-rate-two-peaks is Tier-1 sourced but documents hold-queue abandonment (waiting for a human agent), a mechanism distinct from AI-response abandonment — no primary source was found confirming the vendor's specific claim that voice-AI response delays above ~1200ms cause hangups. That figure stays [unverified-quant — needs primary]. Multi-vendor hop-cost (~250ms) also remains unaddressed by this search. Status stays open, narrowed to: (a) a primary source for AI-response-latency-to-abandonment specifically, and (b) an independent measurement of multi-vendor STT→LLM→TTS hop cost.