{
  "id": 13211257,
  "title": "1 Year of Building a \"Bring Your Own API Key\" Transcription App — and Why It's Now an Automation Hub",
  "url": "https://urgent.news/2026/10/09/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-10-09T20:48:10.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/whisperdirect/1-year-of-building-a-bring-your-own-api-key-transcription-app-and-why-its-now-an-automation-hub-57jk"
  },
  "original_language": "en",
  "account": "A year ago, a solo iOS developer released a transcription app named WhisperDirect. The app allowed users to use their own OpenAI API key directly with the Whisper API, essentially serving as a vessel for using OpenAI via individual API keys. This approach eliminated the need for a backend to maintain and reduced middleman markups, as users paid OpenAI directly for their usage.\n\nWhisperDirect began its development with a simple idea: use the user's OpenAI API key for Whisper API directly. Initially, the app was a one-time purchase with no subscription. Users were billed by OpenAI for their API usage, and the app's backend was not maintained by the developer.\n\nThe app started with a very basic transcription feature, utilizing Whisper large-v3-turbo on a GPU server. However, the usage never warranted running a GPU 24/7, leading to increased bills. Consequently, the developer switched to CPU processing and tested various ASR engines, eventually settling on Apple Speech.\n\nDuring this period, the developer also worked on another app, WhisText, which served as a testing ground for new technologies. Over time, WhisperDirect incorporated features from WhisText, such as speaker diarization, OCR for text from images, waveform display, live preview, and custom summarization models from OpenAI, Gemini, or any OpenAI-compatible endpoint.\n\nThe latest version of WhisperDirect, now in App Store review, consolidates transcription, summarization, and automation features. Users can record audio directly into the app and receive live transcription, summaries, and minutes. Additionally, the app supports importing audio and video files, auto-highlighting text synced with playback, and exporting results in various formats, including subtitles. Users can also choose from multiple LLM models for summarization and post data to various platforms such as Slack using webhooks. The app offers a 7-day free trial and charges users only for API usage.",
  "summary": "Hi, I'm a solo iOS developer. About a year ago, when I first released my transcription app, I wrote this: \"The moment you go subscription, you can't stop.\" \"Even if only one person has ever paid, you have to keep the servers running — even at a loss.\" \"A single traffic spike can degrade your service.\" While every AI app was racing toward monthly subscriptions, I realized that for a solo…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}