Urgent.News

What's breaking now, across thousands of outlets.

AI

How I Built a Real-Time Multilingual AI Voice Tutor for Bharat (And Solved the 55ms Latency Problem)

How I Built a Real-Time Multilingual AI Voice Tutor for Bharat (And Solved the 55ms Latency Problem) 10 Days of Voice AI Challenge — VoiceForBharat Edition Author: Jay | Track: Learning & Literacy (EdTech) Open-Source Repository: github.com/jaysid97/Ten-day-of-voice-challenges-bharat-edition When you build a traditional text chatbot, a 1.5-second API delay feels completely normal. But in…

Over the past 10 days, as part of the #VoiceForBharat AI Challenge, I created Shiksha AI — a multilingual voice tutor with sub-100ms response times, using Murf Falcon TTS, LiveKit Agents, Deepgram Nova-3, and Google Gemini. In this article, I’ll discuss the architecture, challenges, and how you can run the entire open-source system in under two minutes.

Traditional text chatbots often accept a 1.5-second API delay, but voice AI demands near-instantaneous responses. Imagine a student asking "Bhaiya, is quadratic equation ko solve karne ka simple trick kya hai?"—a 1.5-second delay feels like a permanent eternity.

Shiksha AI addresses four major issues in Indian education: cost, language barriers (code-mixing between Hindi and English), fear of being shamed, and the awkwardness of typing equations on a small screen. By providing a patient, language-aware voice tutor, we can democratize personalized learning.

The system architecture is a bi-directional pipeline:

1. Audio from the student’s microphone travels through WebRTC or SIP.

2. Speech is transcribed using Deepgram Nova-3 Multi-Language STT, recognizing English, Hindi, and Hinglish.

3. The text transcript moves to the LLM (Google Gemini 1.5 Flash) for reasoning and tool execution.

4. The LLM outputs spoken response text, which is then sent to Murf Falcon Streaming TTS for conversion to high-quality audio (55ms latency).

5. The audio is streamed via LiveKit Agent Transport to the learner.

Key engineering breakthroughs:

1. Using Murf Falcon’s streaming API, Murf instantly synthesizes audio as soon as Gemini generates the first tokens, reducing latency to under 100ms from the moment the student stops speaking.

2. Shiksha AI retains learner context (like "Welcome back Ramesh!") while ensuring consent for storing voice data in SQLite.

3. Multilingual support is achieved with Deepgram Nova-3, handling real-time speech across languages.

4. LiveKit provides the telephony backbone, handling WebRTC audio streaming, voice activity detection, and SIP telephony.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Sydney: Building a Hinglish AI/ML Mentor That Actually Talks — 10 Days of Voice Agents

One day seating in the on the TL;DR Over 10 days, as part of Murf AI's 10 Days of Voice Agents — Voice for Bharat Challenge 2026 , I built Sydney — a Hinglish-speaking AI/ML learning companion that…

  • Sydney AI mentor teaches RAG, backpropagation, embeddings, agent architectures
  • Sydney conversational AI in Hinglish, Hindi-English mix
  • Sydney features include memory, practice exercises, human escalation

More from Saturday 15 August →