Urgent.News

What's breaking now, across thousands of outlets.

AI

This AI research company wants to put 1,000 African languages into AI

African Languages Lab has spent years building data for underrepresented African languages. Now, its Mansa AI platform brings 30 languages into production, with plans to reach 1,000.

This AI research company wants to put 1,000 African languages into AI

For years, Issaka encountered a persistent issue while developing African language technology - scarce and low-quality data. To address this, the AI research and deployment firm he founded in 2020, African Languages Lab, shifted focus to data collection. Today, the company boasts the largest repository of African-language data, encompassing over 70 languages including Amharic, Hausa, Zulu, Twi, Igbo, and Yoruba.

On Tuesday, African Languages Lab unveiled Mansa, a multilingual and multimodal AI platform featuring approximately 30 languages in production, accessible via web, mobile, and API for developers and businesses.

Mansa's datasets comprise more than 100 billion curated tokens and over 19,000 hours of speech recordings validated by language experts. Africa hosts over 2,000 languages, yet fewer than 5% possess the resources necessary for natural language processing. Consequently, most African languages remain inadequately represented in digital datasets, hindering AI's ability to comprehend them effectively.

Issaka highlighted the language modeling challenges, stating that models often struggle to comprehend African languages in full conversations. Additionally, data collection expenses vary significantly between languages. For instance, generating a token in Yoruba can be four times more costly than in English. Safety tuning, crucial for preventing harmful model responses, is typically stronger in languages with abundant training data, leaving under-resourced African languages with weaker safeguards.

Subsequently, researchers experimenting with harmful prompts face up to ten times higher likelihood of encountering erroneous responses in these under-resourced languages.

African Languages Lab has been amassing this data for nearly a decade, employing freely available datasets and direct community collaboration. They seek contributions from individuals eager to provide language data, earning compensation for their efforts. The company's platform, All Voices, facilitates direct data exchange between African languages, eliminating the need for a bridge language. Contributors can earn income for supplying or validating data.

The challenge lies in uneven data distribution. Languages such as Yoruba, Swahili, and Zulu boast more extensive data repositories than smaller, low-resource languages. Most collected data is text-based; however, the company has increasingly invested in speech data as well, recognizing the importance of reaching the most people who require access to such technologies, particularly those who cannot read or write.

To mitigate these issues, Issaka established All Voices six years ago, enabling contributors to collect and validate data directly between African languages, circumventing the need for a bridge language. Contributors can receive financial incentives for supplying or validating data. Notably, All Voices receives two to three weekly inquiries from individuals interested in contributing data in underrepresented languages. Issaka emphasized the community's shared commitment to building with this technology.

African Languages Lab currently supports over 70 African languages, but Mansa's production capabilities are limited to approximately 30 languages. This decision reflects the company's dedication to ensuring high-quality support and validity for the languages they deploy. The remaining languages are not yet prepared for public use.

Written by urgent.news from TechCabal's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techcabal.com →

More in AI

More from Tuesday 8 September →