Urgent.News

What's breaking now, across thousands of outlets.

AI

Whistle: Speech to Text in 16.9 MB

Whistle is an open speech recognition model that operates on the same CPU engine as Needle, allowing it to run on devices such as mobiles, wearables, robots, smart home systems, automotive systems, and microcontrollers. The model, which weighs in at a compact 16.9 MB, runs on the CPU without any dependencies and integrates seamlessly with Needle's C++ engine within the same container.

After receiving audio, the system processes it through a front end which frames 16 kHz mono audio into 80 log-mel bins, band-limited between 250-3500 Hz. Thirty seconds of audio translates to 3,000 frames, and a convolutional stem of 128 channels and a kernel of 9 halves this count to 375 frames, processed one per 80 ms. Subsequent stages run at this rate, with every embedding stage returning a single row per frame.

The encoder comprises eight Simple Attention blocks, including four mHC residual lanes and a Monarch Hadamard MLP, mirroring Needle's structure but with a different layer count. Unlike traditional models, Whistle's attention mechanism is not causal; each frame attends to frames later in time, such as a frame at 3 seconds attending to a frame at 12 seconds.

The decoder employs eight Laddered Simple Attention blocks, each with 512 width, 8 query heads paired with 2 KV heads, and 48-dimensional queries and keys, 64-dimensional values, and a 3-tap causal convolution on Q, K, and V. These projections are computed once when the clip arrives and held for the entire decoding process, enabling efficient five beam scoring based on length-normalized log probability.

Additionally, keyword biasing uses an Aho-Corasick automaton to enhance the log probability of specified phrases during search.

The transcript is limited to 320 tokens, with language detection encompassing eight thousand text pieces and seven language tokens, which are emitted directly as tokens rather than returned separately. The ladder allows for per-depth training, with the encoder remaining unaltered across all depths. The engine measures audio loudness before decoding, ceasing if the audio falls below a threshold.

Whistle outperforms Whisper base on several benchmark tests including LibriSpeech test-clean and test-other, SPGISpeech, Earnings-22, and the FLEURS average, while being narrowly behind Whisper base on TED-LIUM, AMI, and the MLS average, despite its significantly smaller size. This efficiency makes Whistle a promising tool for resource-constrained environments.

Written by urgent.news from Hacker News's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at cactuscompute.com →

More in AI

More from Thursday 8 October →