Urgent.News

What's breaking now, across thousands of outlets.

AI

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles

Mega AI Battle: Benchmarking 6 Top LLMs with Advanced Bangla Logic Riddles ๐ŸŽฏ Hi everyone! I am thrilled to share my project for the Kaggle Benchmarking Challenge . Instead of using standard English datasets, I created a custom evaluation benchmark consisting of highly complex, linguistically trapped Bangla logic riddles to test the actual reasoning capabilities of 6 world-class AI models: Googleโ€ฆ

A groundbreaking AI competition took place, pitting six top language models against a series of challenging Bangla logic riddles. The evaluation focused solely on the models' reasoning capabilities, rather than their text generation skills, using a custom dataset comprised of complex Bangla riddles. The riddles were designed with non-linear geometric and linguistic traps to test the models' true reasoning abilities.

The first riddle, known as the Circular Spatial Trap, involved a table and chairs arranged in a circle. Google Gemini was the only model to correctly answer the question of how many chairs were on the table - 12. Other models, including ChatGPT, Claude, and Grok, failed to provide the correct answer.

The second riddle, called the Linguistic Semantic Trap, introduced a chicken and rabbit problem. The twist was that the riddle explicitly stated to multiply the number of heads by the number of legs, rather than adding them. Every single AI model failed to recognize the trick and instead calculated the riddle using addition, proving that LLMs still struggle with contextual semantics in non-English languages.

Overall, Google Gemini emerged as the top performer, correctly solving the circle spatial trap but falling victim to semantic traps. Claude and Grok both scored 3/5, while ElevenLabs, ChatGPT, and Blink all scored 2/5. The benchmark demonstrates that while modern LLMs excel at text generation, they still face significant challenges when it comes to specialized local language processing and handling complex logic traps.

Written by urgent.news from Dev.to's reporting โ€” not their text. Machine-written โ€” may contain errors; check the original before relying on it.

Read the original at dev.to โ†’

More in AI

Wednesday assorted links

1. Palo Alto Networks (cybersecurity firm, check out YTD). 2. OAI doing math again. Quasi-Riemann! And just one metric of import. 3. The AI agents pitching literary magazines. 4.

I asked an LLM which listings to skip. It kept getting the numbers wrong.

I waste a stupid amount of time reading pages just to find the one line that rules them out. A job post where the budget is hidden at the very bottom.

  • User created AI tool to check listing accuracy
  • Tool splits tasks between AI and human for numeric validation
  • Extension displays listings in green, amber, or red based on checks

๐ŸŒฟ FrostBite: Offline AI Garden Scout

๐ŸŒฟ FrostBite: Offline Open-Weight Garden Scout This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass ๐Ÿ“– What I Built FrostBite is a lightweight, zero-cloud CLI toolโ€ฆ

  • FrostBite is a CLI tool for gardeners to disconnect from screens.
  • Tool runs offline using local USDA data and smollm2:1.7b model.
  • Benefits include privacy, accessibility, customization, and resilience.

PocketPause: one local AI card, then outside

This is my submission for Week 1: Touch Grass . What I built PocketPause is a small local app with one job: give you one outdoor observation, then get out of the way. Choose roughly 5, 10 or 15 minutes and a broad setting such as a street, terrace, courtyard or campus. A local open-weight model picks a few observation cues.

  • PocketPause is a local AI card app for outdoor observations
  • Users choose time (5-15 mins) and location for a single card
  • App runs locally, no account or GPS needed, MIT-licensed

More from Wednesday 7 October โ†’