A modular offline multilingual speech-to-speech translation system with confidence-aware ASR routing and language-aware TTS
Scientific Reports, Published online: 22 August 2026; doi:10.1038/s41598-026-67788-0 A modular offline multilingual speech-to-speech translation system with confidence-aware ASR routing and language-aware TTS
Offline speech-to-speech translation offers privacy and connectivity benefits for constrained applications, but local deployment demands an optimal balance of recognition accuracy, language coverage, and latency. The authors introduce a modular, fully offline, turn-based system that integrates language identification, confidence-aware automatic speech recognition routing, multilingual machine translation, and language-aware speech synthesis.
The system tested with Hindi, Telugu, Tamil, Bengali, French, German, Arabic, and Japanese as source/target languages, using FLEURS, FLORES-200, and CVSS benchmarks. A routing threshold of 0.70 resulted in 97.65% routing accuracy, significantly lowering recognition errors in four Indic languages compared to Whisper-only recognition.
The study demonstrates that clean-versus-cascaded translation experiments show recognition error propagation in machine translation. End-to-end ASR-BLEU scores range from 31.4 to 39.2 across five CVSS-supported source-to-English directions, while eight-language speech-synthesis evaluation yields mean opinion scores between 3.55 and 4.12 and back-ASR WER ranging from 6.8 to 10.8%.
The modular system's mean model-only CPU latency is 25.41 seconds, with translation and speech synthesis consuming the majority of processing time. The results illustrate the accuracy-precision-latency trade-offs in offline multilingual speech translation on local CPU execution.
Written by urgent.news from Scientific Reports's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.