Urgent.News

What's breaking now, across thousands of outlets.

AI

Building African-language AI is easier than finding the data

The shortage of African-language text data threatens to limit the development of AI tools that reflect the continent’s linguistic diversity.

Building African-language AI is easier than finding the data

Vambo AI, a Johannesburg-based AI model builder, has developed a 1.5-billion-parameter model called MORENA that covers 12 African languages. The company struggled to find sufficient real African-language text data, leading to the use of synthetic data. This data shortage has the potential to limit the development of AI tools that accurately reflect the linguistic diversity of the continent.

MORENA was trained on 251.7 billion tokens, including 65.6 billion African-language tokens. The model performs strongly on African-language text modelling, with a score of 1.408 bits per byte across the 12 languages, the lowest among 26 models tested. Vambo overcame the lack of compute power by leveraging the Leonardo supercomputer provided by the Italian consortium CINECA through the UNDP AI Hub for Sustainable Development.

The company positions MORENA as infrastructure rather than a finished consumer product, with weights available under an Apache 2.0 licence.

Written by urgent.news from TechCabal's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at techcabal.com →

More in AI

I asked 13 AI models what they look like. None chose a human body.

I asked AI models a simple question: how do you imagine yourself? Not necessarily in human form, in any form at all. If you could be seen, what would you look like? I expected variety.

  • None of 13 AI models chose a human body as their self-image
  • Most models described a glowing network or lattice in darkness
  • Three newest models (Gemini 3.8 Flash, Grok, ChatGPT) depicted a translucent glass polyhedron

More from Monday 28 September →