How AI Voice Agents Actually Work
AI phone agents went from obviously robotic to occasionally indistinguishable from a person in about two years. The reason is not one breakthrough. It is that three separate components got good at the same time, and the engineering problem of connecting them got solved well enough. Here is what is actually happening during a call, and why some systems feel natural and others do not. The pipeline…
Artificial intelligence phone agents have improved dramatically over the past two years, moving from obvious robotic responses to ones that are sometimes indistinguishable from those of a human. This progress was not the result of a single breakthrough, but rather the simultaneous advancement of three key components and the effective engineering of their integration.
During a call, a voice agent operates through a pipeline of three models and telephony. The first step involves converting the caller's audio into text using speech recognition. Modern systems do this in real time, continuously transcribing the conversation rather than waiting for the caller to finish speaking. Next, the language model takes the transcribed text, along with the conversation history and any business knowledge provided to the agent, to determine what to say and sometimes what actions to take.
The final step involves converting the model's response back into speech and sending it back down the line, also in a streaming manner so that speech begins before the full response is generated. Latency, or the delay between the caller finishing speaking and the agent beginning to speak, is the critical factor determining whether the call feels natural.
Human conversation typically features gaps of 200 to 300 milliseconds, and a call that feels natural should not exceed approximately 1.5 seconds. Factors affecting latency include audio transport, transcription, the language model's initial token production, speech synthesis, and the return of audio to the speaker. Each component must be fast, and the architecture should be designed to overlap them instead of running them sequentially.
Streaming at every stage is crucial for achieving this. Choosing the right models involves a trade-off between accuracy and speed. Larger models generally produce better responses but take longer to process, while smaller models are faster but may not be as accurate. Many production systems employ a smaller model for the main conversation and switch to a larger one when the request demands it.
Turn-taking, or determining when the caller has finished speaking, can be challenging and often exposes the limitations of a system. Simple methods use voice activity detection, setting a threshold for silence to signal the end of a turn. However, these approaches can make agents appear impatient or laggy. More sophisticated systems use semantic endpointing, analyzing the content of the utterance to determine if it sounds complete.
For instance, "My name is Sarah and my number is is clearly unfinished even after a two-second pause," while "My name is Sarah, and my number is" is recognized as complete. Another important capability is barge-in, which allows the caller to interrupt the agent mid-sentence, causing the agent to stop talking and listen. A system that interrupts the caller mid-sentence feels robotic and is an easy way to evaluate a product's quality.
Good voice synthesis goes beyond just producing correct words; it also handles prosody, such as rhythm, emphasis, and intonation. Older synthesis systems produced correct words but with a flat delivery, while current models can convey emphasis, natural pauses, and rising intonation on questions. Specific telltale signs of poor voice synthesis include the incorrect reading of proper nouns, street names, unusual surnames, and numbers with improper grouping.
Testing a voice agent's synthesis quality with local addresses, particularly outside of English-speaking regions, is essential. A system trained primarily on English may struggle with French street names and regional pronunciation, highlighting the importance of testing in the intended deployment environment. Voice agents that only converse are limited in their usefulness.
More valuable agents can perform lookups and take actions based on the conversation. Two mechanisms enable this: retrieval and tool calling. Retrieval involves providing the agent with access to reference materials, such as business hours, services, pricing, and policies, allowing the agent to pull relevant information into the conversation to maintain accuracy and stay on topic.
Tool calling enables the model to invoke functions, such as checking calendars, creating tickets, looking up orders, sending messages, and transferring calls. This is where most practical value lies, and it is crucial to inquire about this functionality when evaluating a product, as a demo of natural conversation does not necessarily indicate the agent's ability to perform actual tasks.
Configuration of a voice agent can be a significant product decision. Older systems required users to write prompts, which works for developers but fails for businesses that need voice agents. A better approach is to ask structured questions about the business, such as activities, services, hours, tone, escalation rules, and use the answers to build the configuration.
Mirage Cloud's receptionist agent, for example, follows this approach with a guided setup and a browser test before attaching a phone number, allowing businesses to identify potential issues themselves rather than through customer complaints. Before purchasing a voice agent, it is essential to test several aspects to ensure the product meets your needs.
Interrupting the agent mid-sentence should reveal whether it stops and listens or talks over you. Pausing mid-sentence should show if the agent waits or jumps in. Providing the agent with local addresses and unusual surnames can test its ability to handle complex inputs. Finally, asking the agent to perform specific tasks, such as booking, looking up information, or transferring calls, will help determine its conversation quality and task completion capabilities.
Conversation quality and task completion are distinct capabilities, and evaluating both is crucial for selecting the right voice agent for your business.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.