Tone

How a voice AI pipeline actually works, end to end

What happens between a caller finishing a sentence and an agent answering it: voice activity detection, streaming speech recognition, the language model, speech synthesis, and the four places it usually breaks.

The Tone team ·

A voice agent looks like one thing to the caller and is five things underneath. Understanding where the boundaries are is what makes the difference between debugging the pipeline and guessing at it.

This describes the component architecture — separate speech recognition, language model and speech synthesis — which is what Tone (usetone.ai) runs and what most production phone agents run today.

One turn, five stages

A caller says something. Here is everything that happens before they hear a reply.

1. Voice activity detection decides when they stopped

Before anything can be transcribed, something has to decide that the caller has finished a sentence rather than paused mid-thought.

This is harder than a silence timer, and getting it wrong is felt immediately in both directions. Too eager and the agent interrupts someone who was thinking. Too patient and every exchange has an awkward hole in it. Real implementations use a minimum speech duration before declaring speech started, and a silence tail — commonly around 600ms — before declaring it ended.

That silence tail is worth noticing: it is latency the caller pays on every turn, it happens before any model is involved, and no amount of GPU makes it smaller.

2. Speech recognition turns audio into text

Audio streams to a recognition service over a websocket while the caller is still talking. A good one emits interim transcripts — best guesses, revised as more audio arrives — and then a final one when the utterance closes.

Interims matter for more than a live caption. They let you start the next stage speculatively, before the final transcript exists, which is one of the few ways to remove time from the critical path rather than just moving it.

3. The language model decides what to say

The transcript, the conversation history, the system prompt and any retrieved knowledge go to a language model, which streams tokens back.

The number that matters here is time to first token — how long before the model emits anything at all. Not total generation time: you do not wait for the whole reply. As soon as you have a complete clause you can start speaking it, and the model keeps generating while the caller listens.

TTFT is typically the largest single stage in the pipeline, and it is genuine server-side generation time. Moving your infrastructure closer to the model does not reduce it.

4. Speech synthesis turns text into audio

Clauses go to a text-to-speech service as they become available, and audio comes back in chunks that go straight onto the call.

This is where an unintuitive constraint lives: TTS is usually billed per character at submission, and it is typically the largest provider cost in the whole pipeline — around 95% of provider spend in ours. Which means the agent talking less is the main cost lever, and that every time a caller interrupts a reply you already submitted, you paid for audio nobody heard.

5. The transport carries it

On a phone call, a carrier bridges the caller to your media endpoint and streams audio both ways. This is also where the constraint that shapes everything else lives: telephony audio is narrowband and lossy, which is why recognition accuracy on real calls is not what a demo with a laptop microphone suggests.

The four things that usually break

Barge-in that is too sensitive

The caller must be able to interrupt. But a naive implementation — "any sound above the noise floor stops the agent" — is cut short by a cough, a door, or a notification chime.

The defence is layered: a minimum speech duration before anything counts, then withholding the interruption until either a transcript with actual words arrives or speech continues past a confirmation threshold. A blip that produces neither leaves the reply playing.

There is a nastier version. Near-silence makes recognition models hallucinate — "Thank you." is the canonical output for line noise. A wordy hallucination looks exactly like a real interruption, cuts the agent off mid-greeting, and the transcript afterwards looks perfectly clean. Only the recording shows it happened.

Language detection that outruns the voice

Recognition models generally understand more languages than synthesis models can speak. Detect a language your voice model cannot produce and the reply fails or arrives in the wrong language.

The fix is to clamp a detection to the languages the agent is configured to speak, rather than adopting whatever came back. An empty configured list should mean "no opinion, adopt the detection" rather than "English only" — the difference decides whether a code-switching caller is handled or contradicted.

For Indian languages there is a further trap: a short answer is often transcribed in Latin script and tagged as English. A caller's "हाँ" can arrive as Ha, tagged en-IN, two characters long. A model that has not been told that Latin "Ha" is हाँ will answer it with filler instead of advancing the conversation. This never shows up in browser-microphone testing, where the audio is cleaner and the tester is usually speaking English — which is exactly why these bugs appear "only on real calls".

Answering machines

A voicemail greeting is speech, so the agent talks to it. It then bills a full conversation against a mailbox, and the campaign records a successful contact.

Detection is the easy half. The hard half is that an answering machine cannot be interrupted — it keeps talking regardless — so every caller-protection mechanism you built works against you. Machine audio arriving during your voicemail message reads as a barge-in and cancels the hang-up you just armed. The whole voicemail sequence has to be guarded as one unit, not check by check.

Lazy initialisation hiding in turn one

Anything resolved on first use — credentials, connection pools, model warm-up — lands inside the caller's first question.

We measured a multi-second penalty on turn one from exactly this: minting a cloud credential lazily instead of at boot. Turn one is the turn that decides whether someone stays on the line, and it is the one least likely to show up in an average.

Why a component pipeline rather than one model

End-to-end speech-to-speech models are real and getting better, and they remove the inter-stage handoffs entirely.

The reason to keep components separate is control and evidence. With a pipeline you can see each stage's timing independently, swap a provider without rewriting the agent, constrain what the agent is allowed to say at the text layer, and get a transcript as a by-product rather than as a second inference.

For regulated calling that last point is not a nicety. If you have to be able to show months later what was said on a specific call, a pipeline that produced the text on the way past is a fundamentally different position from one that has to re-derive it.

Where to go next

If you want the numbers rather than the architecture, the stage-by-stage latency breakdown has what each of these costs in milliseconds, measured in production in Mumbai.

Common questions

What are the components of a voice AI agent?
Five, in a loop per turn: voice activity detection to know when the caller started and stopped speaking, streaming speech recognition to turn their audio into text, a language model to decide the reply, speech synthesis to turn that reply into audio, and a telephony transport carrying audio both ways. Each is a separate network service, and the caller experiences the sum.
What is barge-in in a voice agent?
Barge-in is the caller interrupting the agent mid-sentence and the agent stopping to listen. Without it the agent talks over the caller, which makes a conversation unusable. Implementing it well is harder than it sounds, because the system has to distinguish a real interruption from a cough, a notification sound, background speech, or the agent's own audio echoing back through a speakerphone.
Why does a voice agent sometimes answer in the wrong language?
Usually because speech recognition understands more languages than speech synthesis can speak. If recognition detects a language the voice model cannot produce, the reply either fails or comes back in an unintended language. The fix is to clamp the detected language to the set the agent is actually configured to speak, rather than adopting whatever was detected.
Is a speech-to-speech model better than a pipeline of separate components?
It depends what you need to control. An end-to-end model removes inter-stage latency and can sound more natural, but a component pipeline lets you see and fix each stage independently, swap a provider, enforce what the agent may say, and get a text transcript as a by-product rather than as an extra inference. For regulated calling, where you must be able to prove what was said, the transparency usually wins.

In the docs

Related

Build this on Tone

Tone (usetone.ai) provisions +91 phone numbers over an API and runs the TRAI compliance those calls are subject to. Test calls are free on a sandbox key.