Making voice-agent failures reproducible with Python
A voice agent can answer correctly and still make a conversation frustrating. It can wait too long before replying, keep speaking after an interruption, or confirm the wrong appointment time after a correction. Those failures need more than a transcript review. Why build a tool for this? I'm building a voice agent as a hobby project: real-time STT, LLM and TTS, the whole loop. Open-source voice…
Voice agents can provide correct answers yet frustrate users with issues like delayed replies, speaking over interruptions, or confirming incorrect appointment times after corrections. Building a tool to evaluate such failures is essential. The author created a Python-based voice-agent evaluation framework called voice-evals to address these concerns.
The project offers a CLI and library that allows users to replay recorded calls or scripted live conversations against any WebSocket-based voice agent. It provides deterministic scoring, per-stage latency budgets, barge-in stop times, and CI gating capabilities. Users can input recorded calls or scripted calls through a compatible WebSocket transport and export the recording for later replay.
The CLI tool accepts various evaluation parameters, such as maximum WER, minimum task completion, maximum p95 latency, and outputs the results in JSON format. If any evaluation thresholds are breached, the tool exits with a code of 1. The framework also includes a probe that can simulate conversations with an agent, helping with regression testing.
The framework supports various input sources, including recordings, scripted calls, and different TTS providers. It can handle mono PCM audio at 16 or 24 kHz and includes a reference agent for testing socket communication. Users can configure different WebSocket providers, such as ElevenLabs, and integrate their own services.
The framework records detailed evidence, including manifest files, configuration, event logs, and audio files. It allows for reproducibility of evaluation results by running the exported calls against the same thresholds. This ensures that the evaluation of a recorded conversation can be replicated and compared against future versions of the voice agent.
The framework also differentiates between per-call latency measures and per-session timing values. It provides turn-level latency percentiles for individual turns and session-level latency percentiles for the overall call. These metrics serve different purposes and should not be used interchangeably.
In summary, voice-evals is a comprehensive framework for evaluating voice agents, covering aspects like transcription accuracy, per-stage latency, barge-in handling, and scripted conversation replay. By providing deterministic scoring, per-stage latency budgets, and robust evidence collection, the framework enables developers to identify and address failures in voice agent implementations.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.