{
  "id": 13720067,
  "title": "Speaker Labels Burned a Full Hour of Compute and Returned Nothing. One Voice Print Per Phrase Fixed It",
  "url": "https://urgent.news/2026/10/11/speaker-labels-burned-a-full-hour-of-compute-and-returned-nothing-one",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-11T13:12:20.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/nikita_iakovlev_415524c19/speaker-labels-burned-a-full-hour-of-compute-and-returned-nothing-one-voice-print-per-phrase-fixed-4f0f"
  },
  "original_language": "en",
  "account": "In the latest development, a transcriber built for an Apify actor failed to deliver results in a full hour of compute despite being fed with a 62-minute lecture. A full hour of compute was spent, yet no output was generated. The error message stated that the actor run had reached the 3600-second timeout and was aborted. The issue stemmed from the actor's design, which computed speaker embeddings every 10 seconds at a one-second step, resulting in each second of audio being embedded roughly ten times. This design proved inefficient for processing longer audio files on a 2 GB memory limit, causing it to run slower than Whisper, the transcription model itself. The transcriber attempted to address the issue by making two changes: first, it stopped computing embeddings for every 10-second window and instead focused on embedding speech itself; second, it changed the approach to segmentation, identifying stretches of speech and cutting the audio at pauses using two small ONNX models. The updated actor processed the same file with the same memory limit, completing the transcription in 8365 words at an estimated 437 seconds. The change resulted in a significant reduction in compute usage, down to 0.24 units from 2.0 units, leading to a successful transcription of the entire lecture. However, a potential limitation was identified in the form of incomplete speaker identification when two people speak over each other continuously without a pause.",
  "summary": "I run a visa agency in Bali. The part of my work that has nothing to do with visas is building Actors on Apify, and one of them is a transcriber: you give it an audio or video URL, it runs Whisper large v3 and gives back text with timecodes. Transcription was the easy half. The feature that nearly did not ship was the one that sounds like a checkbox — \"label who is speaking.\" The run that…",
  "key_points": [
    "Speaker embeddings computed 10 times per second, inefficient for long audio",
    "Updated actor focused on embedding speech, segmented audio at pauses",
    "Compute usage reduced to 0.24 units, successful transcription achieved"
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}