{
  "id": 13015999,
  "title": "Whistle: A 16.9 MB Speech-to-Text Model That Runs on Any CPU (Hands-On Test)",
  "url": "https://urgent.news/2026/10/09/whistle-a-16-9-mb-speech-to-text-model-that-runs-on-any-cpu-hands-on",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-09T03:08:33.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/jamilxt/whistle-a-169-mb-speech-to-text-model-that-runs-on-any-cpu-hands-on-test-3175"
  },
  "original_language": "en",
  "account": "Whistle is an open-source speech recognition model developed by Cactus Compute and released on October 2, 2026. This speech-to-text model, available as a single 16.9 MB file, operates on CPU with no dependencies and boasts accuracy numbers that surpass Whisper base. Despite its small size, which is smaller than most app icon assets, Whistle delivers impressive performance, transcribing a public speech recording correctly and providing word-level timestamps within 10 seconds on a conventional CPU.\n\nThe Whistle model supports seven languages - English, German, French, Spanish, Italian, Dutch, and Polish - and automatically detects the language unless it is manually specified. For each word, the model provides a start time, end time, and probability score, along with speech embeddings for search and matching purposes.\n\nWhistle's architecture consists of an audio encoder with eight attention blocks, a small decoder, and gated cross attention. A key engineering choice in the development of Whistle is its use of quantization, making it a 2 to 4 bit model and enabling it to fit within the 16.9 MB file size. Comparatively, Whisper base, which operates in floating-point format, weighs in at 145.3 MB.\n\nFor testing, the model was installed on a plain Linux VPS without any GPU or specialized tooling, utilizing real audio. The installation process involved a few simple commands, and the model's size was verified to be exactly 16,919,407 bytes. The transcription of a public domain speech recording from Wikimedia Commons yielded a correct and punctuated transcription, along with accurate timestamps and confidence levels for each word.\n\nWhile vendor-reported accuracy numbers and latency figures are presented in the test, it is essential to note that these are separate from the actual performance observed during testing. The Whistle model demonstrates the potential for on-device speech recognition, eliminating the need for network calls, reducing marginal costs to zero, and addressing privacy concerns by keeping audio data within the device.",
  "summary": "Say you set up your agent stack to accept voice notes, the way most voice-first projects start. The plan is predictable: record audio, hit a transcription API, get text back. At $0.006 per minute it sounds free, until you remember that a voice-first agent loop transcribes everything it hears, including the garbage. Then a post hits the Hacker News front page and stops you: a speech-to-text model…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}