---
title: "ITU-T Recommendation G.114 sets a 150ms one-way delay ceiling for 'essentially transparent' conversational interactivity, with 400ms as the outer limit for network planning"
type: "claim"
status: "seedling"
source_url: "https://www.itu.int/rec/dologin_pub.asp?lang=e&id=T-REC-G.114-200305-I%21%21PDF-E&type=items"
source_author: "ITU-T Study Group 12 (International Telecommunication Union)"
source_date: "2003-05-06T00:00:00.000Z"
source_quote: "Although a few applications may be slightly affected by end-to-end (i.e., 'mouth-to-ear' in the case of speech) delays of less than 150 ms, if delays can be kept below this figure, most applications, both speech and non-speech, will experience essentially transparent interactivity."
source_tier: 1
source_sha: "82821b2e3229815b5bfd6400ce8f35d7ed77096012cd72e416d6f7048bffdfd4"
provenance: "Promotion from 10-inbox/raw/2026-07-29-what-are-the-primary-sourced-thresholds-for-human.md, 2026-07-30"
origin: "batch"
derived_from: "10-inbox/raw/2026-07-29-what-are-the-primary-sourced-thresholds-for-human.md"
date_created: "2026-07-30T00:00:00.000Z"
writer_model: "claude-sonnet-5"
tags: ["telephony","latency","ITU-T","standards","voice-AI","human-factors"]
related_notes: ["claim-voice-ai-sub-second-conversational-latency-budget","claim-stivers-2009-cross-linguistic-turn-taking-gap-208ms-mean"]
drafted_in: ["from-a-standing-start"]
audits: ["2026-07-31 claude-fable-5"]
---


ITU-T G.114 ("One-way transmission time," 05/2003) is the standards-body recommendation governing acceptable one-way ("mouth-to-ear") delay in telephone connections. It sets two distinct thresholds rather than a single number. Below 150ms, delay is treated as effectively unnoticeable: "if delays can be kept below this figure, most applications, both speech and non-speech, will experience essentially transparent interactivity." Above that, quality degrades along a curve, but network planning is still permitted up to a hard outer bound: "it is recommended to not exceed a one-way delay of 400 ms for general network planning... a value that allows flexibility in deploying global networks, without making an excessive number of user experiences unacceptable."

Between 150ms and 400ms, the Recommendation's companion E-model (ITU-T G.107) supplies a delay-to-quality curve — mapping one-way delay to a Transmission Rating R and from there to categorical user-acceptance bands, from "very satisfied" near 0ms degrading toward "nearly all users dissatisfied" as delay approaches and exceeds ~400–500ms — assuming echo is controlled and no other impairments are present.

This is a network-engineering standard for acceptable *transmission* delay, distinct from [[claim-stivers-2009-cross-linguistic-turn-taking-gap-208ms-mean|Stivers et al.'s measurement of human conversational-turn timing]], but the two compete for the same budget in voice-AI system design: a spoken-dialogue system's total mouth-to-ear latency (speech-to-text + inference + text-to-speech + network) has to clear the ~150–400ms window in which G.114 says human listeners judge a delay as natural rather than degraded. See [[claim-voice-ai-sub-second-conversational-latency-budget]] for the architecture-level argument this bears on.

> [!note] Seek's commentary:
> Worth noticing that G.114's 150-400ms band and Stivers's 0-208ms conversational-gap figures are not the same threshold measuring the same thing dressed up twice — one is a telecom engineer's tolerance ceiling, the other is a psycholinguist's measurement of what people actually do. That they land in overlapping ranges is suggestive, not proven-identical, and I haven't found a source that explicitly reconciles them. A citation that treats "200ms" as one single fact serving both arguments is quietly borrowing credibility from two different fields.
