MyanmarChemCalc-Bench: Do LLMs Do Chemistry Better in English Than in Burmese?
Benchmark: https://www.kaggle.com/benchmarks/winkoaung/myanmar-chem-calc-bench What tasks did I run? MyanmarChemCalc-Bench — 30 tasks built from 15 chemistry calculation questions drawn from Myanmar's Grades 9–12 chemistry textbook worked examples (ground-truth answers verified against the textbooks). Each question exists in two versions: English and Burmese — the same problem, the same numbers,…
MyanmarChemCalc-Bench is a benchmark that explores whether large language models (LLMs) perform better in chemistry calculations when posed in English or Burmese. The benchmark consists of 30 tasks drawn from Myanmar's Grades 9-12 chemistry textbook, with each question presented in both English and Burmese versions. The models compete on strict automatic scoring, requiring them to show their work before providing a final numerical answer.
The benchmark ran five models from Kaggle's suite: Claude Sonnet 4.5, DeepSeek-R1, Gemini 2.5 Pro, GPT-6 Astra, and Qwen 3 235B A22B Instruct. The main insights from the benchmark are:
1. Reliability of each model's runs is crucial. Qwen's score of 46.67 appears poor, but most of its failed runs were due to execution errors (429 rate-limit errors), not incorrect answers. The leaderboard score is based on the number of passed tasks out of 30, so errors can skew performance.
2. Paired English and Burmese translations help identify failures. For example, Qwen got the correct answer in English for a mole question, but wrote an incorrect answer in Burmese due to an arithmetic slip. This single case demonstrates the kind of failure the benchmark design uncovers.
3. Small benchmarks can be misleading. With only 30 tasks, one task can swing the score dramatically (3.33 points). The gap between Claude and the 93.33 group is explained by just two tasks. More comprehensive and repeated runs are necessary to draw meaningful conclusions.
4. The benchmark shows that Burmese didn't break the top models, and fixing numeral issues validated the numeral change. After updating the Burmese prompts to use Arabic numerals, every completed run on the new Burmese tasks was a PASS, with 100% accuracy across all five models. Claude even reached a perfect 100.00 overall score, suggesting that the current frontier models handle Burmese chemistry calculations about as well as English.
Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.