Urgent.News

What's breaking now, across thousands of outlets.

AI

Voxlocal: a minimal voice agent written in Rust

It's essential to understand that voice agents are innovative pieces of technology designed to process natural human-like spoken conversations, completing tasks as requested. Voice agents listen, think, perform actions, and speak, making them fascinating pieces of technology. This article aims to explore the ins and outs of how these voice agents function.

Telephony networking plays a crucial role in voice agents, bridging traditional public telephone networks with cloud-based artificial intelligence applications. This complex process involves handling call lifecycles, bi-directional streaming, transcoding, and low latency. A significant task for this segment is jitter buffering, which addresses dropped or delayed packets, while also ensuring audio upsampling to maintain Speech-to-Text (STT) accuracy.

Voxlocal is a minimal voice agent developed for macOS, written in Rust. It enables users to interact with the voice agent via a command-line interface (CLI) and receive responses to their queries. Created for experimentation purposes, its primary focus is on an automotive-service use case, using English language commands. The program leverages pre-built components, such as Whisper and Piper, to achieve this functionality.

When interacting with a voice agent, it processes audio samples by breaking them down into smaller tasks. It recognizes words from the audio, cleans up the text, determines the relevant information, selects an appropriate action, and converts the response back into speech. In Voxlocal, the microphone captures sound which is measured continuously and transformed into audio samples.

The captured audio undergoes several transformations, including mono conversion and resampling to 16,000 samples per second. This processed audio is then passed to Whisper, a trained neural network designed to identify speech patterns and estimate sequences of words. Voxlocal utilizes greedy decoding to select the most likely next token, which, in this case, is the word "oil" or "change."

Once Whisper provides a response in segments, Voxlocal joins these segments into a single transcript. However, due to factors such as noise, accents, cadence, and timbre, the output may not perfectly match the actual spoken words. To address this, Voxlocal employs an extra stage to clean up the transcribed text, making it more consistent for further processing.

Humans often add filler words, such as "Um," or present peculiar pronunciations, which can hinder the accuracy of speech recognition. Voxlocal employs simple and predictable rules to remove these common filler phrases and whitespace, allowing the next stages to interpret the intended meaning of the user's query. Regular expressions used in the normalizer are reused on each turn, adhering to the program's minimalist approach.

Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at samkhawase.com →

More in AI

Free AI video quotas worth knowing (Oct 2026 caps)

Short list of free or freemium video-generation allowances that still show up when you need a clip for a demo or placeholder.

  • Runway offers about 125 free credits on the free plan, last checked in October 2026.
  • Replicate provides around $5 in free credits for new accounts, last checked in October 2026.

FieldQuest: one local AI quest to get you outside

This is a submission for the Hacktoberfest Open-Source AI Challenge: Week 1: Touch Grass . What I Built FieldQuest creates one small outdoor quest from two simple choices: how much time you have and…

  • FieldQuest generates outdoor quests based on user inputs of available time and mood.
  • App operates offline, utilizing Hugging Face Transformers.js and SmolLM2-360M-Instruct model.
  • No location data requested, no server employed, all user inputs stored in page memory.

More from Saturday 10 October →