An offline voice agent for my terminal in 52 MB (Whistle + Needle)
Two of the biggest Hacker News threads this week were about Whistle, a 16.9 MB speech-to-text model, and DeepSeek V4.1 Flash. I wanted to see what you can build with the first one, so I put a voice front end on my terminal, with the second as a fallback. Say "show me the last fifteen minutes of logs for the checkout service" and it prints kubectl logs deploy/checkout --since=15m . Speech…
Two Hacker News threads this week discussed Whistle, a 16.9 MB speech-to-text model, and DeepSeek V4.1 Flash. The author sought to demonstrate what could be achieved with Whistle by attaching a voice interface to their terminal, using DeepSeek as a backup. The command "show me the last fifteen minutes of logs for the checkout service" translates to "kubectl logs deploy/checkout --since=15m".
The audio remains within the machine; speech recognition and tool selection are processed by the CPU. The local model, Whistle, handles multiple languages. Needle, also from the package, is a small model for tool selection. Combined, both models occupy 52 MB of disk space. The package pip install cactus-needle downloads the models upon first use.
To enable microphone input, the [mic] extra is added, and pip install openai is used for the fallback. The script features six tools and a confirmation gate, with a cloud fallback when Needle cannot produce usable output. To use the script, provide a WAV file or allow it to record four seconds from the microphone by default. The script logs each command instead of executing it.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.