{
  "id": 9576967,
  "title": "Pipecat Benchmarked 23 Real-Time STT Models for Voice Agents. There Isn’t One Winner.",
  "url": "https://urgent.news/2026/09/24/pipecat-benchmarked-23-real-time-stt-models-for-voice-agents-there",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-24T15:43:05.000Z",
  "source": {
    "name": "HackerNoon",
    "slug": "hackernoon",
    "url": "https://hackernoon.com/pipecat-benchmarked-23-real-time-stt-models-for-voice-agents-there-isnt-one-winner?source=rss"
  },
  "original_language": "en",
  "account": "Pipecat ran an open-source benchmark with 1,000 real utterances, measuring both latency and semantic accuracy. They released the code and data so others could replicate the results. The benchmark did not declare a clear winner but created a curve of models that trade off speed for meaning. Models close to the Pareto frontier, which represents the best trade-off between latency and accuracy, are defensible options. The decision of which model to choose depends on how much latency and semantic error your specific application can tolerate. Semantic Word Error Rate (WER) measures whether the meaning was preserved, not just if the words matched. It's a better proxy for real-world interaction success than raw WER, which weights all errors equally. However, the benchmark's choice of audio, transcripts, and annotation method limits its applicability to all use cases. The fastest model may not always provide the most accurate or reliable results for your specific use case.",
  "summary": "Pipecat benchmarked 23 real-time STT models for voice agents. See how latency and accuracy compare, and why the best model depends on your use case.",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}