Voxlocal: a minimal voice agent written in Rust
It's essential to understand that voice agents are innovative pieces of technology designed to process natural human-like spoken conversations, completing tasks as requested. Voice agents listen, think, perform actions, and speak, making them fascinating pieces of technology. This article aims to explore the ins and outs of how these voice agents function.
Telephony networking plays a crucial role in voice agents, bridging traditional public telephone networks with cloud-based artificial intelligence applications. This complex process involves handling call lifecycles, bi-directional streaming, transcoding, and low latency. A significant task for this segment is jitter buffering, which addresses dropped or delayed packets, while also ensuring audio upsampling to maintain Speech-to-Text (STT) accuracy.
Voxlocal is a minimal voice agent developed for macOS, written in Rust. It enables users to interact with the voice agent via a command-line interface (CLI) and receive responses to their queries. Created for experimentation purposes, its primary focus is on an automotive-service use case, using English language commands. The program leverages pre-built components, such as Whisper and Piper, to achieve this functionality.
When interacting with a voice agent, it processes audio samples by breaking them down into smaller tasks. It recognizes words from the audio, cleans up the text, determines the relevant information, selects an appropriate action, and converts the response back into speech. In Voxlocal, the microphone captures sound which is measured continuously and transformed into audio samples.
The captured audio undergoes several transformations, including mono conversion and resampling to 16,000 samples per second. This processed audio is then passed to Whisper, a trained neural network designed to identify speech patterns and estimate sequences of words. Voxlocal utilizes greedy decoding to select the most likely next token, which, in this case, is the word "oil" or "change."
Once Whisper provides a response in segments, Voxlocal joins these segments into a single transcript. However, due to factors such as noise, accents, cadence, and timbre, the output may not perfectly match the actual spoken words. To address this, Voxlocal employs an extra stage to clean up the transcribed text, making it more consistent for further processing.
Humans often add filler words, such as "Um," or present peculiar pronunciations, which can hinder the accuracy of speech recognition. Voxlocal employs simple and predictable rules to remove these common filler phrases and whitespace, allowing the next stages to interpret the intended meaning of the user's query. Regular expressions used in the normalizer are reused on each turn, adhering to the program's minimalist approach.
Written by urgent.news from Lobsters's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.