Urgent.News

What's breaking now, across thousands of outlets.

AI

Making voice-agent failures reproducible with Python

A voice agent can answer correctly and still make a conversation frustrating. It can wait too long before replying, keep speaking after an interruption, or confirm the wrong appointment time after a correction. Those failures need more than a transcript review. Why build a tool for this? I'm building a voice agent as a hobby project: real-time STT, LLM and TTS, the whole loop. Open-source voice…

Voice agents can provide correct answers yet frustrate users with issues like delayed replies, speaking over interruptions, or confirming incorrect appointment times after corrections. Building a tool to evaluate such failures is essential. The author created a Python-based voice-agent evaluation framework called voice-evals to address these concerns.

The project offers a CLI and library that allows users to replay recorded calls or scripted live conversations against any WebSocket-based voice agent. It provides deterministic scoring, per-stage latency budgets, barge-in stop times, and CI gating capabilities. Users can input recorded calls or scripted calls through a compatible WebSocket transport and export the recording for later replay.

The CLI tool accepts various evaluation parameters, such as maximum WER, minimum task completion, maximum p95 latency, and outputs the results in JSON format. If any evaluation thresholds are breached, the tool exits with a code of 1. The framework also includes a probe that can simulate conversations with an agent, helping with regression testing.

The framework supports various input sources, including recordings, scripted calls, and different TTS providers. It can handle mono PCM audio at 16 or 24 kHz and includes a reference agent for testing socket communication. Users can configure different WebSocket providers, such as ElevenLabs, and integrate their own services.

The framework records detailed evidence, including manifest files, configuration, event logs, and audio files. It allows for reproducibility of evaluation results by running the exported calls against the same thresholds. This ensures that the evaluation of a recorded conversation can be replicated and compared against future versions of the voice agent.

The framework also differentiates between per-call latency measures and per-session timing values. It provides turn-level latency percentiles for individual turns and session-level latency percentiles for the overall call. These metrics serve different purposes and should not be used interchangeably.

In summary, voice-evals is a comprehensive framework for evaluating voice agents, covering aspects like transcription accuracy, per-stage latency, barge-in handling, and scripted conversation replay. By providing deterministic scoring, per-stage latency budgets, and robust evidence collection, the framework enables developers to identify and address failures in voice agent implementations.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

I told an AI "put me in the rain with a bee circling my head" — and heard it orbit my ears

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend Spatial Audio Sandbox — describe a scene, hear it in 3D Flutter + Rust HRTF engine · Gemma-3-1B on Render · ElevenLabs…

  • Developer built Spatial Audio Sandbox for textual-to-audio experience
  • Gemma model generates screenplay JSON with sources and positions
  • Rust HRTF mixer renders sounds in real-time with head tracking

I Made ScamLens So My Family Could Check Suspicious Messages Without Sending Them to the Cloud

This is my submission for the Hacktoberfest 2026 Weekend Challenge: Build for a Friend . What I Built My family gets scam messages all the time. Fake bank alerts. KYC requests.

  • ScamLens app analyzes messages and screenshots locally without cloud AI
  • Demonstrated with phishing, social engineering, and legitimate OTP scenarios
  • Fix added semantic fields to distinguish message types for accurate classification

More from Monday 5 October →