Urgent.News

What's breaking now, across thousands of outlets.

AI

DarijaBench: Do AI Models Actually Understand Moroccan Darija?

DarijaBench: Do AI Models Actually Understand Moroccan Darija? As a Moroccan student, I use AI assistants every day. They are brilliant in English and French — but I kept noticing something: ask them something in Darija (Moroccan Arabic dialect, spoken by 35+ million people), and the confident answers start wobbling. So I decided to stop guessing and start measuring. I built DarijaBench, a…

A Moroccan student named DarijaBench has created a benchmark to measure how well AI models understand Moroccan Darija, a dialect spoken by over 35 million people. This 60-item test evaluates the models on three practical skills: translating from Darija to French, sentiment analysis, and answering questions about Morocco. The benchmark was built using the Kaggle Benchmarks SDK and ran on four AI models: Gemini 3.7 Flash, Claude Sonnet 5, Claude Opus 4.7, and GPT-5.6 Luna.

The results showed a three-way tie at 0.95 for the top models, which was unexpected because Darija is often considered a weak spot for AI. Sentiment analysis proved to be the easiest task, with Gemini even scoring a perfect 20/20. However, translation from Darija to French was the hardest, with a 85% success rate for Gemini. This difficulty comes from Darija's unique features, such as French loanwords and idioms that don't have direct equivalents in French.

GPT-5.6 Luna scored the lowest at 0.90, which, despite being close to the other models, still led to around 3 more failures on the test. This discrepancy is likely due to the limited training data on Moroccan specifics. The creator of the benchmark cautions that keyword-based grading can be brittle, and more challenging items, like idioms and code-switching, are needed for a more accurate comparison.

The benchmark is now available for others to use and improve upon. If you speak a low-resource dialect, you can build your own version using the Kaggle Benchmarks SDK. This democratization of benchmarking allows for more accurate assessments of AI's ability to understand diverse languages and dialects.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Thursday 1 October →