Urgent.News

What's breaking now, across thousands of outlets.

AI

Running Local AI Speech-to-Text in Browser CPU with WebAssembly: Zero Server Uploads

Beaming confidential client interviews, confidential board meetings, or unreleased podcast recordings to remote cloud speech-to-text APIs is a massive privacy risk. Every major "AI Transcription" startup asks you to upload your audio files to their cloud S3 buckets. Once uploaded, your voice data sits on remote servers subject to data leaks or model retraining. With modern WebAssembly SIMD and…

Uploading sensitive audio files to remote cloud speech-to-text APIs poses a significant privacy concern. Leading AI transcription companies require users to send their voice data to their servers, leaving it vulnerable to breaches or model retraining. However, with the advent of WebAssembly SIMD and ONNX Runtime Web, it's now possible to execute quantized Whisper transformer models entirely within a user's browser, thanks to SolveMyMedia Transcribe.

SolveMyMedia Transcribe leverages Transformers.js to decode audio into 16kHz float buffers in browser memory and processes them directly through a local model runtime. To get started, import the pipeline function from @xenova/transformers and initialize the Whisper Tiny English model using the Xenova/whisper-tiny.en endpoint, specifying the webgpu device for computation.

Then, use the transcribeLocalAudio function to pass an audioBlob, which is converted to an arrayBuffer and passed to the transcriber. The function returns the transcription text, ensuring that no data leaves the user's machine.

When comparing the privacy and cost implications, the Cloud Speech API, provided by major services like OpenAI or Rev, stores audio files on external cloud servers, incurring a cost of $0.006 per minute. In contrast, SolveMyMedia Local AI Transcribe keeps all data within the client's RAM, incurring zero costs and ensuring complete privacy. Additionally, the local solution supports offline operation without internet dependency, unlike the cloud-based API, which requires an internet connection.

Try out SolveMyMedia Transcribe yourself at https://solvemymedia.com/transcribe and share your experiences with client-side AI inference in production in the comments section!

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Is a README for Humans or for LLMs?

I read the README of one of my own repositories and stopped at the tenth sentence. I had written it with my README skill, and the pre-publication check had said it was fine to publish.

  • Separate README sections for humans and LLMs improve clarity
  • Human-readable section contains concise, essential information
  • LLM-readable section includes additional facts for AI assistants

Claude Haiku 5.5 Pricing: The 100K-Token Rule for Agents

Claude Haiku 5.5 costs $0.10 per million input tokens up to 100K tokens of prompt, and $0.50 per million above it, a fivefold cliff that decides your agent bill.

  • Claude Haiku 5.5 introduces 100K-token cliff
  • Pricing $0.10 per million input up to 100K tokens
  • Output costs $0.50 per million below limit

Kazakhstan seeks to build an AI-powered economy and manufacturing sector

Kazakhstan wants artificial intelligence to become more than a digital tool: the country is seeking to make it a new engine of economic growth, from universities and factories to energy and…

  • Kazakhstan aims to become an AI-powered economy and manufacturing sector.
  • AI should augment human capabilities and improve production efficiency.
  • Kazakhstan is developing all five layers of the AI economy.

More from Thursday 8 October →