The Open ASR Leaderboard Adds Its First Global South Language
Voice Arena and Hugging Face have teamed up to launch an open ASR evaluation for Hindi and Indian English. This evaluation introduces two new sets: Monsoon en-IN and Monsoon hi-IN. Hindi, spoken by over half a billion people, becomes the first Indic language on a multilingual leaderboard that currently only includes European languages.
Each set is made publicly available for self-scoring, while a private split is withheld to prevent benchmark-specific optimization. The sets vary across nine axes, including geography, age, gender, vocabulary, devices, acoustic environments, speech type, speech rate, and multiple valid transcripts for the same audio. This allows test sets to expose failure modes that are specific to certain populations.
The data for Monsoon was sourced from unscripted dual-channel spontaneous conversations, with each clip carrying information about the speaker's occupation, education, marital status, income band, handset brand, current city, and years in the current district. While the sets are small in terms of hours, they are large in terms of the number of speakers, ensuring diversity and avoiding overfitting to specific voices, regions, or device models.
Indian English is represented across six zones, with each zone well represented in the datasets. Unlike most benchmarks, Monsoon data comes from unscripted conversations and includes detailed demographic information, allowing for fine-grained analysis of ASR performance across various factors. This open ASR evaluation on a public leaderboard enables a more comprehensive analysis of ASR capabilities and errors, helping to identify and address disparities in automated speech recognition across different populations and regions.
Written by urgent.news from Hugging Face's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.