{
  "id": 10710730,
  "title": "Why we don’t proxy live audio WebSockets through Vercel (and what we do instead)",
  "url": "https://urgent.news/2026/09/29/why-we-dont-proxy-live-audio-websockets-through-vercel-and-what-we-do",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-29T14:33:15.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/yvone_f8de85837dee3e4cd0f/why-we-dont-proxy-live-audio-websockets-through-vercel-and-what-we-do-instead-7go"
  },
  "original_language": "en",
  "account": "Transcribing live audio on Vercel proved challenging because the platform expects a request/response model. Live audio streams generate a continuous flow of bidirectional bytes that cannot be effectively handled by serverless functions. These limitations include cold starts mid-session, hard execution time limits, awkward billing for idle keep-alives, and potential security risks from long-lived connections and API key exposure. To address these issues, the team established a rule: do not proxy long-lived audio or high-frequency events through the app server. Instead, live audio should reside on an edge provider or a dedicated long-running Node process.\n\nThe recommended approach involves the following four jobs, each with clear responsibilities:\n1. Your API: Handles authentication, credit pre-checks, job creation, and generates a scoped, short-TTL token.\n2. Provider: Browsers establish a WebSocket connection with the provider directly, receiving partial transcripts.\n3. Your API (optional): Performs mid-session metering to warn users or cut off when the wallet reaches zero.\n4. Your API: Finalizes the transcript, charges the last partial minute, and marks the job as complete.\n\nTo maintain fair billing, a simple floor/ceil pattern is used for metering. Users must pre-check by having at least one minute of credits. Metering posts to your API occur every few seconds, charging by the floor minutes so far (zero until 60 seconds). Finalize ensures the last partial minute is billed as well. This billing method is idempotent and performed via HTTPS posts from the client, keeping the WebSocket on the provider's edge. Your database serves as the source of truth for tracking billable minutes consumed by each job.\n\nWhile provider-direct connections suffice for speech-to-text and many chat streams, dedicated Node processes (like Fly, Railway, or a VM) are recommended for scenarios requiring:\n- Custom audio mixing or VAD before the vendor processes the bytes\n- Multi-party fan-in that no single provider session can handle\n- A protocol the browser cannot safely speak with a short token\n\nBy dividing responsibilities between your API (control plane) and the ASR vendor (data plane), you create a more efficient, scalable, and secure live transcription solution on Vercel. This approach has been successfully implemented for live transcription and live translation (captions plus sentence-level translation) on TranscribeChirp. When wiring a similar architecture on Next.js, focus on the mental model rather than the specific brand name.",
  "summary": "When we shipped live captions for TranscribeChirp , the first instinct was tempting: Put a WebSocket on the Next.js API route, pipe mic audio through Vercel, fan out to Deepgram. That path fails on serverless. Here is the split we ended up with — and why it is a better default for any “live while the user is still talking” feature on Vercel. The constraint Vercel (and similar platforms) want…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}