{
  "id": 3474960,
  "title": "AI models flub these intelligence tests. Can you fare any better?",
  "url": "https://urgent.news/2026/08/26/ai-models-flub-these-intelligence-tests-can-you-fare-any-better",
  "topic": "ai",
  "section": "AI",
  "published": "2026-08-26T09:00:00.000Z",
  "source": {
    "name": "MIT Technology Review",
    "slug": "mit-technology-review",
    "url": "https://www.technologyreview.com/2026/08/26/1141952/puzzles-ai-models-flub-these-tests/"
  },
  "original_language": "en",
  "account": "In the realm of artificial intelligence, puzzles and games have long served as benchmarks to gauge the progress of AI models. From the early days of machine learning, developers have relied on gaming challenges to assess how far AI models have come. Arthur Samuel's 1959 article on an algorithm that learned to play checkers marked the beginning of this trend. Chess and the Chinese board game Go have also become iconic AI test beds.\n\nDespite the strides made in AI capabilities, puzzles continue to reveal the technology's strengths and weaknesses. While models have shown significant improvement, they still struggle with certain types of puzzles. Subtle changes in classic riddles, as well as visual puzzles, often trip them up. This article provides a series of puzzles designed to test your wits against AI, highlighting the differences between human and machine cognition.\n\nOne area where humans excel is spatial reasoning. Mental rotation problems, which involve determining whether different images represent the same objects from various angles, are a prime example. While language models can analyze visual inputs, they struggle to manipulate 3D objects the way humans do. In a study by researchers from Google and the University of Illinois Urbana-Champaign, models only figured out 18% of the New York Times Connections puzzles, while some could solve them near-perfectly by early 2025.\n\nAnother domain where humans outperform AI is memory and adaptability. LLMs have vast memories and can recite many facts faithfully, making them excel at trivia. However, when a puzzle closely resembles one the model has encountered during training, it may fail to distinguish key differences and provide memorized responses instead. This phenomenon was observed in a 2024 study involving Knights and Knaves puzzles, where models struggled with slight variations of classic puzzles they had seen before.\n\nAbstract and visual reasoning also present challenges for AI models. ARC-AGI, a benchmark requiring models to infer abstract rules from examples, has seen some improvement in recent years. However, even top-tier models can be stumped by specific puzzles. For instance, a problem presented here relies on a transformation rule that models have difficulty identifying, highlighting the limitations of their visual reasoning abilities.\n\nInterestingly, humans are not immune to cognitive traps as well. Psychologists have developed problem suites that exploit errors in intuitive math and phrasing, often leading to non-deliberate answers. These mental shortcuts, known as heuristics, can lead humans to give knee-jerk responses, while AI models may respond more deliberatively. The Lightning Round section below challenges you to answer questions quickly, emphasizing the importance of critical thinking and careful analysis in the face of cognitive biases.",
  "summary": "Puzzles and games have been central to AI development since the very beginning. Just as we humans like to test our smarts with crosswords or logic puzzles, developers can test how far models have advanced with a gaming gauntlet. The term “machine learning” was popularized in a 1959 article by the IBM computer scientist Arthur…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}