Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles
Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles ๐ฏ Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge . Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Googleโฆ
A groundbreaking AI competition took place, pitting six top language models against a series of challenging Bangla logic riddles. The evaluation focused solely on the models' reasoning capabilities, rather than their text generation skills, using a custom dataset comprised of complex Bangla riddles. The riddles were designed with non-linear geometric and linguistic traps to test the models' true reasoning abilities.
The first riddle, known as the Circular Spatial Trap, involved a table and chairs arranged in a circle. Google Gemini was the only model to correctly answer the question of how many chairs were on the table - 12. Other models, including ChatGPT, Claude, and Grok, failed to provide the correct answer.
The second riddle, called the Linguistic Semantic Trap, introduced a chicken and rabbit problem. The twist was that the riddle explicitly stated to multiply the number of heads by the number of legs, rather than adding them. Every single AI model failed to recognize the trick and instead calculated the riddle using addition, proving that LLMs still struggle with contextual semantics in non-English languages.
Overall, Google Gemini emerged as the top performer, correctly solving the circle spatial trap but falling victim to semantic traps. Claude and Grok both scored 3/5, while ElevenLabs, ChatGPT, and Blink all scored 2/5. The benchmark demonstrates that while modern LLMs excel at text generation, they still face significant challenges when it comes to specialized local language processing and handling complex logic traps.
Written by urgent.news from Dev.to's reporting โ not their text. Machine-written โ may contain errors; check the original before relying on it.