Urgent.News

What's breaking now, across thousands of outlets.

Tech

How to transcribe audio and video files to text with one API call (MP3, MP4, Google Drive, Dropbox)

Speech-to-text models are very good now. The annoying part is everything around them: recordings that are too big to upload, video files, and files that live behind a Google Drive or Dropbox share link. This post covers the do-it-yourself route first, then a one-call alternative I built for myself. The do-it-yourself route Short audio file If you have a short MP3 and Python, open-source Whisper…

Transcribing audio and video files to text can now be done with a single API call, regardless of whether the files are MP3, MP4, hosted on Google Drive or Dropbox. The process begins with speech-to-text models which are now very good, but the tricky part is dealing with recordings that are too large to upload or files that reside behind a share link. The post covers both a do-it-yourself route and a one-call alternative.

For the DIY route, short audio files can simply be processed using open-source Whisper. For video files, the audio track must first be extracted with ffmpeg, and the file then compressed to ensure compatibility with speech models that work on 16 kHz mono. Long recordings and hosted APIs offer faster transcription but have limits on upload size. Long recordings must be split into chunks, transcribed individually, and then timestamps adjusted and the text combined.

Share links require additional steps, as they open a preview page rather than the file itself. Dropbox links need to have dl=1 appended, while Google Drive share links need the file ID extracted and used in a specific format. These are small steps, but they can be numerous and complex to manage.

The one-call route simplifies this with a tool called Audio & Video to Text on Apify. This tool handles extraction, splitting, retries, and timestamps. Whisper large-v3 is used for transcription. The response is a JSON object per file containing the transcription text, duration, and an SRT file. This can be easily integrated into Python scripts for batch processing.

In practice, a 50-minute MP3 shared via Dropbox was transcribed in about 40 seconds at a cost of $0.15, with a rate of $0.003 per audio minute. The tool does not support YouTube or TikTok links, and it requires files rather than video pages. Speaker labels are not provided, and the audio is sent to a hosted provider, which may be a concern for sensitive recordings. However, for a few short files and scenarios where privacy matters most, running Whisper on your own machine remains a viable option.

The tool was built by the author of this article and this article was drafted with AI assistance and checked by the author.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in Tech

Nightly Pipeline Errors: Express API Production Logs for Tenant Run Reconstruction

TL;DR: For a B2B SaaS nightly data pipeline, choose log management around incident reconstruction, not the prettiest dashboard.

  • Centralize log management for incident reconstruction in B2B SaaS pipeline.
  • Emit structured JSON with stable correlation fields like runid, tenantid, jobname.
  • Exclude secrets and direct identifiers from event emission to protect data privacy.

More from Monday 5 October →