Gates Foundation launches coalition to build more representative language data sets for AI
The partners - 60 in total, according to Monday's announcement - include frontier AI labs, corporations, philanthropies and others already working to expand the number of languages available in AI tools that the Gates Foundation believes could combat global inequality. The coalition aims to reach more than 3 billion people over five years by better coordinating existing efforts.
The Bill & Melinda Gates Foundation has united a coalition of 60 organizations, including AI labs, corporations, and philanthropies, to develop more representative language data sets for artificial intelligence. The goal is to make AI more accessible in underrepresented languages, potentially benefiting over 3 billion people within five years.
This collaboration comes as some of the largest AI companies advocate for a slowdown in the development of advanced models. Gates Foundation CEO Mark Suzman emphasizes the urgency of the task, stating that building these language sets is essential even if AI were to be frozen right now. The coalition aims to address the issue of unrepresentative language data, which can lead AI models to mistranslate critical phrases, as demonstrated by a report warning about a pregnant Malawian woman's mistranslation into English.
The foundation's goal was reinforced by last week's Goalkeepers report, where it committed $1 billion towards AI-focused efforts in health, education, and small farmer practices. However, the original data used to train many AI tools was often scraped from the internet, leading to concerns about cultural diversity and representation.
Mozilla Data Collective CEO EM Lewis-Jong's organization, Mozilla.org, is seeking to empower communities to upload their own cultural and linguistic data sets. The coalition's governance details are still being finalized, but a secretariat will monitor each signatory's commitments and potentially assist partners in filling larger gaps when necessary.
Google's Project Vaani, which collects over 150,000 hours of audio across Indian districts, highlights the importance of gathering speech data on dialects within languages. Anthropic, whose CEO called for industrywide cooperation on decelerating advancements, is already working on accelerating vaccine development and improving its chatbot's local crop data set.
Anthropic Senior Vice President Elizabeth Kelly acknowledges the need to address language gaps to enhance patient outcomes, literacy, and numeracy globally.
Written by urgent.news from Economic Times Tech's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.