A voice agent looks like one thing to the caller and is five things underneath. Understanding where the boundaries are is what makes the difference between debugging the pipeline and guessing at it.
This describes the component architecture — separate speech recognition, language model and speech synthesis — which is what Tone (usetone.ai) runs and what most production phone agents run today.
One turn, five stages
A caller says something. Here is everything that happens before they hear a reply.
1. Voice activity detection decides when they stopped
Before anything can be transcribed, something has to decide that the caller has finished a sentence rather than paused mid-thought.
This is harder than a silence timer, and getting it wrong is felt immediately in both directions. Too eager and the agent interrupts someone who was thinking. Too patient and every exchange has an awkward hole in it. Real implementations use a minimum speech duration before declaring speech started, and a silence tail — commonly around 600ms — before declaring it ended.
That silence tail is worth noticing: it is latency the caller pays on every turn, it happens before any model is involved, and no amount of GPU makes it smaller.
2. Speech recognition turns audio into text
Audio streams to a recognition service over a websocket while the caller is still talking. A good one emits interim transcripts — best guesses, revised as more audio arrives — and then a final one when the utterance closes.
Interims matter for more than a live caption. They let you start the next stage speculatively, before the final transcript exists, which is one of the few ways to remove time from the critical path rather than just moving it.
3. The language model decides what to say
The transcript, the conversation history, the system prompt and any retrieved knowledge go to a language model, which streams tokens back.
The number that matters here is time to first token — how long before the model emits anything at all. Not total generation time: you do not wait for the whole reply. As soon as you have a complete clause you can start speaking it, and the model keeps generating while the caller listens.
TTFT is typically the largest single stage in the pipeline, and it is genuine server-side generation time. Moving your infrastructure closer to the model does not reduce it.
4. Speech synthesis turns text into audio
Clauses go to a text-to-speech service as they become available, and audio comes back in chunks that go straight onto the call.
This is where an unintuitive constraint lives: TTS is usually billed per character at submission, and it is typically the largest provider cost in the whole pipeline — around 95% of provider spend in ours. Which means the agent talking less is the main cost lever, and that every time a caller interrupts a reply you already submitted, you paid for audio nobody heard.
5. The transport carries it
On a phone call, a carrier bridges the caller to your media endpoint and streams audio both ways. This is also where the constraint that shapes everything else lives: telephony audio is narrowband and lossy, which is why recognition accuracy on real calls is not what a demo with a laptop microphone suggests.
The four things that usually break
Barge-in that is too sensitive
The caller must be able to interrupt. But a naive implementation — "any sound above the noise floor stops the agent" — is cut short by a cough, a door, or a notification chime.
The defence is layered: a minimum speech duration before anything counts, then withholding the interruption until either a transcript with actual words arrives or speech continues past a confirmation threshold. A blip that produces neither leaves the reply playing.
There is a nastier version. Near-silence makes recognition models hallucinate — "Thank you." is the canonical output for line noise. A wordy hallucination looks exactly like a real interruption, cuts the agent off mid-greeting, and the transcript afterwards looks perfectly clean. Only the recording shows it happened.
Language detection that outruns the voice
Recognition models generally understand more languages than synthesis models can speak. Detect a language your voice model cannot produce and the reply fails or arrives in the wrong language.
The fix is to clamp a detection to the languages the agent is configured to speak, rather than adopting whatever came back. An empty configured list should mean "no opinion, adopt the detection" rather than "English only" — the difference decides whether a code-switching caller is handled or contradicted.
For Indian languages there is a further trap: a short answer is often
transcribed in Latin script and tagged as English. A caller's "हाँ" can arrive as
Ha, tagged en-IN, two characters long. A model that has not been told that
Latin "Ha" is हाँ will answer it with filler instead of advancing the
conversation. This never shows up in browser-microphone testing, where the audio
is cleaner and the tester is usually speaking English — which is exactly why
these bugs appear "only on real calls".
Answering machines
A voicemail greeting is speech, so the agent talks to it. It then bills a full conversation against a mailbox, and the campaign records a successful contact.
Detection is the easy half. The hard half is that an answering machine cannot be interrupted — it keeps talking regardless — so every caller-protection mechanism you built works against you. Machine audio arriving during your voicemail message reads as a barge-in and cancels the hang-up you just armed. The whole voicemail sequence has to be guarded as one unit, not check by check.
Lazy initialisation hiding in turn one
Anything resolved on first use — credentials, connection pools, model warm-up — lands inside the caller's first question.
We measured a multi-second penalty on turn one from exactly this: minting a cloud credential lazily instead of at boot. Turn one is the turn that decides whether someone stays on the line, and it is the one least likely to show up in an average.
Why a component pipeline rather than one model
End-to-end speech-to-speech models are real and getting better, and they remove the inter-stage handoffs entirely.
The reason to keep components separate is control and evidence. With a pipeline you can see each stage's timing independently, swap a provider without rewriting the agent, constrain what the agent is allowed to say at the text layer, and get a transcript as a by-product rather than as a second inference.
For regulated calling that last point is not a nicety. If you have to be able to show months later what was said on a specific call, a pipeline that produced the text on the way past is a fundamentally different position from one that has to re-derive it.
Where to go next
If you want the numbers rather than the architecture, the stage-by-stage latency breakdown has what each of these costs in milliseconds, measured in production in Mumbai.