Urgent.News

What's breaking now, across thousands of outlets.

AI

Whistle: Speech to Text in 16.9 MB

Article URL: https://cactuscompute.com/blog/whistle Comments URL: https://news.ycombinator.com/item?id=50008427 Points: 304 # Comments: 73

Whistle is a speech recognition model designed for mobiles, wearables, robots, smart homes, automotive systems, and microcontrollers. It weighs in at a compact 16.9 MB file, runs on the CPU with no dependencies, and integrates seamlessly with Needle, a separate open speech recognition model.

This 16.9 MB model performs three key functions: the front end, the encoder, and the decoder. Audio input is processed by framing at a 25 ms window and a 10 ms hop, resulting in 80 log-mel bins that are band-limited to 250-3500 Hz and normalized per channel. Thirty seconds of audio yields 3,000 frames. The convolutional stem processes these frames, reducing the count three times to leave 375 frames at one per 80 ms.

The encoder comprises eight Simple Attention blocks, with four mHC residual lanes and a Monarch Hadamard MLP replacing the feed-forward network, mirroring Needle's structure. Unlike Needle, the attention in Whistle isn't causal; it allows a frame at 3 seconds to attend to a frame at 12 seconds. The decoder features eight Laddered Simple Attention blocks with a width of 512, 8 query heads to 2 KV heads, 48-dimensional queries and keys, 64-dimensional values, and 64-dimensional values. These layers include a unique addition per layer.

The speech-specific part adds one enhancement per decoder layer, using a gated cross attention mechanism. Each decoder layer reads the encoder through this mechanism, with a gate learned per layer and K and V values taken from the clip itself. These projections are computed once when the clip arrives and held for the entire decode process.

Whistle limits its transcript to five short transcript caches, not five passes over the audio, and outputs five beams that are scored by length-normalised log probability. Keyword biasing allows users to boost the log probability of specific phrases during the search. The transcript is capped at 320 tokens, with a vocabulary of 8,192 text pieces and seven language tokens, enabling the model to identify and emit the detected language as a token rather than returning it out of band.

Written by urgent.news from Hacker News Best's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at cactuscompute.com →

More in AI

More from Thursday 8 October →