Most published voice-AI latency numbers are measured from the United States, against endpoints hosted in the United States, by vendors whose infrastructure is in the United States. If your callers are in India, those numbers do not describe your call.
This is what a production voice agent actually costs in AWS ap-south-1
(Mumbai), stage by stage, measured on a real nine-turn phone call rather than
modelled. Tone (usetone.ai) is a +91 telephony API, so every number here comes
from the pipeline we run for customer calls.
The stages of one conversational turn
A voice agent is not one request. It is a chain, and the caller experiences the whole chain. Instrumenting the total tells you nothing about which part to fix, so we timestamp seven points:
| Marker | What just happened |
|---|---|
T0 | The caller stopped speaking — detected by voice activity detection, not by silence timeout |
T1 | Speech recognition returned the final transcript |
TL | The language model's response headers arrived — connected, no tokens yet |
T2 | The language model's first token |
TF | The first complete clause was handed to speech synthesis |
T3 | Speech synthesis returned the first audio chunk of the real reply |
T4 | The first audio byte reached the caller |
Two derived numbers matter, and they are not the same thing:
answer = T3 - T0— when the substantive reply starts. This is the honest engineering number.perceived = T4 - T0— when the caller hears anything. This is what the conversation feels like.
The measurements
Nine-turn call, all stages in milliseconds. The left column is the same code driven from a developer laptop; the right is in-region.
| Stage | From a laptop | In-region (ap-south-1) |
|---|---|---|
asrConnectMs — opening the speech socket | 357–674 | 53 |
asrFinalMs — audio to final transcript | 180–300 | 188–463 |
llmConnectMs — request to first token | 290–1100 | 381–587 |
clauseMs — first token to first speakable clause | 70–380 | 73–436 |
ttsSynthMs — clause to first audio chunk | 220–290 | 199–220 |
answerMs (p50) | 1233 | 1220 |
perceivedMs | 0–3 | 0–1 |
Three things in that table are worth more than the total.
Call setup collapsed, and generation did not
asrConnectMs fell from about 620ms to 53ms. That is what co-location buys,
and it is a real win — but notice that answerMs barely moved, from 1,233 to
1,220. The laptop's round-trip was hiding inside a stage that was mostly not
network.
llmConnectMs only lost the laptop's transport. It never was mostly network:
Vertex flushes its response headers together with the first token, so that stage
is the model's genuine server-side time-to-first-token — around 400ms for Gemini
2.5 Flash — and no amount of proximity removes it.
This is the part most latency advice gets backwards. Moving your infrastructure closer does not speed up generation. It speeds up everything around generation, which in a voice pipeline is four separate socket establishments per turn.
Perceived latency is near zero, and that is not a trick
perceivedMs of 0–1 does not mean the reply arrives instantly. It means the
agent emits a short acknowledgement at T0 while the rest of the pipeline runs
behind it — the same thing a human does when they say "right" before answering.
This is worth being precise about, because it is easy to mistake for a benchmark game. The substantive answer still takes 1,220ms. What changes is that the caller is never sitting in silence wondering whether the line dropped, and silence is what makes people talk over an agent.
Turn one is not like the others
In Vertex mode the application-default credential has to be minted before the first request: read the service-account JSON, sign a JWT, exchange it for a token. Do that lazily and it lands inside the first turn.
Measured, on turn one: llmConnectMs of 2,855ms and 6,169ms, against 285–782
on turns two onward. The caller's very first question — the one that decides
whether they stay on the line — paid a multi-second penalty.
The fix is to mint the token at process boot instead, and re-mint it on a timer so it never expires mid-call. It is a small change and it is the single largest latency win in this entire document.
What this means if you are building one
Measure from end-of-speech, not from your API call. A number that starts when you send the request has already skipped voice activity detection and speech recognition, which together are 240–520ms of the caller's wait.
Instrument stages, not totals. Every optimisation listed above was found by looking at one stage in isolation. A total would have shown 1,233 → 1,220 and suggested nothing was worth doing.
Check what your first turn costs. Lazy initialisation of anything — credentials, connection pools, model warm-up — hides in turn one, which is the turn that matters most and the one least likely to appear in an average.
Region beats model selection, until it doesn't. We benchmarked alternatives:
gemini-2.5-flash at p50 438ms against gemini-3.5-flash at 624ms, with the
newer model also producing roughly twice the reply length — which is more speech
synthesis time and more money on a per-character bill. The faster, shorter
model stayed. Note also that every *-flash-lite variant is global-endpoint
only, with no APAC region at all, so its model card is irrelevant to an Indian
deployment.
Reproducing this
The time-to-first-token benchmark is a command, not a spreadsheet:
pnpm bench:ttft
It drives the real provider including the keep-alive pool, and prints p50/p95 TTFT, first-chunk size and reply length per model. Measure in the region you will deploy in — a ranking taken from a laptop is still a valid ranking, but the absolute numbers are not yours.
Context for these numbers
Independent benchmarks of the major voice-agent platforms report median time-to-first-audio-byte figures of roughly 1,520ms (Bland), 1,558ms (Vapi) and 1,740ms (Retell). Those are measured against US-hosted infrastructure and are honest numbers for that setup.
We are publishing an India-region breakdown because we could not find one, not to claim a win in a benchmark that measures something slightly different. If you are choosing a platform for Indian callers, the question to ask any vendor is not "what is your latency" but "which region do your speech and language models run in, and can you show me the stage breakdown?"