Urgent.News

What's breaking now, across thousands of outlets.

AI

What I Learned Evaluating LLMs Across Four Languages

Most writing about large language models assumes English. The benchmarks are in English, the failure examples are in English, and the quiet implication is that whatever is true in English is true everywhere else, just a little worse. After spending a long stretch evaluating model output across English, Hindi, Tamil, and Marathi, I can say that assumption is wrong in ways that actually matter —…

Many discussions about large language models (LLMs) assume that English is the baseline, and their performance in other languages can be considered slightly worse. However, after evaluating LLMs across English, Hindi, Tamil, and Marathi, this assumption proved to be wrong in significant ways that are crucial for understanding how these models function.

The quality gap among languages is not consistent or where one might expect it to be. While it is tempting to think that English would be the best, and other languages would be consistently worse, the reality is far more complex. A model may excel at conversational Hindi but fail to grasp Tamil's technical vocabulary. Additionally, a model can create fluent Marathi with subtle factual inaccuracies, while producing slightly awkward English that is factually correct.

The issue lies in the interaction between the language and the task rather than the language itself. By evaluating the same task across multiple languages, one can observe these differences that single-language tests often fail to reveal. In English, reviewers have become accustomed to the quirks of model output, making them more discerning even when a confidently wrong answer is presented. In contrast, a language where fluent model output is less common may lead to a higher acceptance of a confident-but-wrong answer.

Errors in multilingual outputs, such as dropping or mangling diacritics, mixing scripts, or incorrectly transliterating, are more apparent in languages with complex scripts like Tamil and Marathi. These errors are often overlooked by monolingual English reviewers and can easily be mistaken for accurate answers by automated exact-match checks. For building evaluation tools for multilingual outputs, it is essential to pay as much attention to script handling and normalization as to semantic correctness.

Code-switching, where speakers frequently mix languages, poses another challenge. While some models handle this naturally, gracefully incorporating English terms within a non-English sentence, others struggle. A single English word can disrupt the grammar of the entire sentence or the model may overcorrect by translating a term that should remain in English. These edge cases are rarely encountered in single-language benchmarks, and they often separate what sounds fluent from what is actually correct.

While the models are improving rapidly, single-language evaluation still provides only a limited understanding of the model's performance across different languages and tasks. The real-world implications of these findings extend beyond Hindi, Tamil, and Marathi. For any product targeting multiple languages, evaluating a single language is insufficient—it measures something different from what is ultimately shipped.

The habit of testing the same task across several languages, and maintaining a skeptical view towards fluency as a proxy for correctness, can significantly improve the evaluation process.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

Run your AI subscription 24 hours a day — use the quota you already pay for

Let me start with a question. Why did I fear development done by artificial intelligence? The answer is plain. AI can build software, and on top of that, it never rests.

  • AI never rests, operating 24/7 without labor laws or human limitations
  • First-mover advantage crucial for AI development, setting standards and pace
  • Unused subscription quota at night presents cost-efficient opportunity for small companies

More from Thursday 3 September →