Tone

Voice AI latency in India: the numbers we measure in ap-south-1

A stage-by-stage latency breakdown of a production voice agent running in Mumbai — where every millisecond of a 1,220ms reply goes, and why in-region matters more than model choice.

The Tone team ·

Most published voice-AI latency numbers are measured from the United States, against endpoints hosted in the United States, by vendors whose infrastructure is in the United States. If your callers are in India, those numbers do not describe your call.

This is what a production voice agent actually costs in AWS ap-south-1 (Mumbai), stage by stage, measured on a real nine-turn phone call rather than modelled. Tone (usetone.ai) is a +91 telephony API, so every number here comes from the pipeline we run for customer calls.

The stages of one conversational turn

A voice agent is not one request. It is a chain, and the caller experiences the whole chain. Instrumenting the total tells you nothing about which part to fix, so we timestamp seven points:

MarkerWhat just happened
T0The caller stopped speaking — detected by voice activity detection, not by silence timeout
T1Speech recognition returned the final transcript
TLThe language model's response headers arrived — connected, no tokens yet
T2The language model's first token
TFThe first complete clause was handed to speech synthesis
T3Speech synthesis returned the first audio chunk of the real reply
T4The first audio byte reached the caller

Two derived numbers matter, and they are not the same thing:

  • answer = T3 - T0 — when the substantive reply starts. This is the honest engineering number.
  • perceived = T4 - T0 — when the caller hears anything. This is what the conversation feels like.

The measurements

Nine-turn call, all stages in milliseconds. The left column is the same code driven from a developer laptop; the right is in-region.

StageFrom a laptopIn-region (ap-south-1)
asrConnectMs — opening the speech socket357–67453
asrFinalMs — audio to final transcript180–300188–463
llmConnectMs — request to first token290–1100381–587
clauseMs — first token to first speakable clause70–38073–436
ttsSynthMs — clause to first audio chunk220–290199–220
answerMs (p50)12331220
perceivedMs0–30–1

Three things in that table are worth more than the total.

Call setup collapsed, and generation did not

asrConnectMs fell from about 620ms to 53ms. That is what co-location buys, and it is a real win — but notice that answerMs barely moved, from 1,233 to 1,220. The laptop's round-trip was hiding inside a stage that was mostly not network.

llmConnectMs only lost the laptop's transport. It never was mostly network: Vertex flushes its response headers together with the first token, so that stage is the model's genuine server-side time-to-first-token — around 400ms for Gemini 2.5 Flash — and no amount of proximity removes it.

This is the part most latency advice gets backwards. Moving your infrastructure closer does not speed up generation. It speeds up everything around generation, which in a voice pipeline is four separate socket establishments per turn.

Perceived latency is near zero, and that is not a trick

perceivedMs of 0–1 does not mean the reply arrives instantly. It means the agent emits a short acknowledgement at T0 while the rest of the pipeline runs behind it — the same thing a human does when they say "right" before answering.

This is worth being precise about, because it is easy to mistake for a benchmark game. The substantive answer still takes 1,220ms. What changes is that the caller is never sitting in silence wondering whether the line dropped, and silence is what makes people talk over an agent.

Turn one is not like the others

In Vertex mode the application-default credential has to be minted before the first request: read the service-account JSON, sign a JWT, exchange it for a token. Do that lazily and it lands inside the first turn.

Measured, on turn one: llmConnectMs of 2,855ms and 6,169ms, against 285–782 on turns two onward. The caller's very first question — the one that decides whether they stay on the line — paid a multi-second penalty.

The fix is to mint the token at process boot instead, and re-mint it on a timer so it never expires mid-call. It is a small change and it is the single largest latency win in this entire document.

What this means if you are building one

Measure from end-of-speech, not from your API call. A number that starts when you send the request has already skipped voice activity detection and speech recognition, which together are 240–520ms of the caller's wait.

Instrument stages, not totals. Every optimisation listed above was found by looking at one stage in isolation. A total would have shown 1,233 → 1,220 and suggested nothing was worth doing.

Check what your first turn costs. Lazy initialisation of anything — credentials, connection pools, model warm-up — hides in turn one, which is the turn that matters most and the one least likely to appear in an average.

Region beats model selection, until it doesn't. We benchmarked alternatives: gemini-2.5-flash at p50 438ms against gemini-3.5-flash at 624ms, with the newer model also producing roughly twice the reply length — which is more speech synthesis time and more money on a per-character bill. The faster, shorter model stayed. Note also that every *-flash-lite variant is global-endpoint only, with no APAC region at all, so its model card is irrelevant to an Indian deployment.

Reproducing this

The time-to-first-token benchmark is a command, not a spreadsheet:

pnpm bench:ttft

It drives the real provider including the keep-alive pool, and prints p50/p95 TTFT, first-chunk size and reply length per model. Measure in the region you will deploy in — a ranking taken from a laptop is still a valid ranking, but the absolute numbers are not yours.

Context for these numbers

Independent benchmarks of the major voice-agent platforms report median time-to-first-audio-byte figures of roughly 1,520ms (Bland), 1,558ms (Vapi) and 1,740ms (Retell). Those are measured against US-hosted infrastructure and are honest numbers for that setup.

We are publishing an India-region breakdown because we could not find one, not to claim a win in a benchmark that measures something slightly different. If you are choosing a platform for Indian callers, the question to ask any vendor is not "what is your latency" but "which region do your speech and language models run in, and can you show me the stage breakdown?"

Common questions

What is a good latency for an AI voice agent?
Under about 1.5 seconds from the caller finishing speaking to the agent starting to speak. Below roughly 800ms a conversation feels natural; above about 2 seconds callers start talking over the agent because they assume the line is dead. Tone (usetone.ai) measures a median of 1,220ms from end-of-speech to the first audio of the substantive reply, running in AWS ap-south-1 (Mumbai).
What is TTFT in a voice AI pipeline?
TTFT is time-to-first-token: how long the language model takes to emit its first token after receiving a prompt. In a voice pipeline it is usually the single largest stage. Measured in-region, Gemini 2.5 Flash returns its first token in roughly 400ms, and no amount of network proximity reduces it because it is server-side generation time, not transport.
Why are voice AI agents slower for Indian callers?
Most voice AI platforms run their speech and language models in US or EU regions. Every turn of the conversation crosses the ocean two or more times, adding roughly 250-400ms of pure transport that never appears in the vendor's own benchmark numbers. Running speech recognition, the language model and speech synthesis inside an Indian region removes that.
How do you measure voice agent latency correctly?
Measure from end-of-speech (when the caller actually stops talking, detected by voice activity detection) to the first audio byte the caller hears. Measuring from the transcript being ready, or from the API call being made, skips real stages and produces numbers that are not what anyone experiences.

In the docs

Related

Build this on Tone

Tone (usetone.ai) provisions +91 phone numbers over an API and runs the TRAI compliance those calls are subject to. Test calls are free on a sandbox key.