Urgent.News

What's breaking now, across thousands of outlets.

AI

Even AI can’t perfectly build Ikea furniture—yet

Building Ikea furniture is not for the faint of heart. So much so that folklore suggests assembling a dresser, a bed, or anything from the Swedish brand can create enough tension it may even destroy a relationship . But beyond the stuff’s being a pain to build, it turns out that the problem-solving and spatial reasoning involved in building Malm dressers or Billy bookshelves are remarkably good…

Even AI can’t perfectly build Ikea furniture—yet

Assembling Ikea furniture is notoriously challenging, often causing stress in relationships, according to folklore. However, the reasoning and problem-solving skills required to build Malm dressers and Billy bookshelves have become useful tools for assessing the intelligence of AI models. Aiden Ament and Greg Burnham, researchers at Epoch AI, a nonprofit AI research organization, recently developed the Furniture Assembly Benchmark to evaluate AI's ability to detect errors in partially assembled furniture.

This benchmark focuses on the models' capacity to recognize mistakes in IKEA furniture builds, which is a real-world application of their problem-solving and spatial reasoning skills.

The researchers began developing this benchmark during a company retreat, eventually settling on the idea of having the AI spot mistakes made during assembly. Three Ikea products with varying levels of complexity were chosen for the experiment: the Ställ shoe rack (easy), the Tonstad bed frame (medium), and the Gullaberg dresser (most complex).

The team built each piece intentionally making mistakes, capturing 60 photos of different assembly stages for each model. They provided the models with assembly instructions, zoom-in tools, and a Python interpreter.

The benchmark was not only about building Ikea furniture but also about measuring how well these models could guide and troubleshoot more complex builds. If AI can reason quickly and accurately in this domain, it could potentially provide reliable guidance for complex physical tasks like repairing cars or household appliances. The benchmark was initially tested on Anthropic’s Claude Opus 4.5, which scored about 28% accuracy. However, within 10 months, OpenAI's GPT-6 Astra model reached 80% accuracy.

Notably, the fastest model, GPT-6 Astra, took around 3 minutes per photo, which is up to 10 times faster than other models. Chinese models, however, lagged behind U.S. models in progress. While the progress in spatial reasoning and troubleshooting is promising, the models still face limitations, particularly in speed and consistent accuracy.

Despite these challenges, experts like Burnham emphasize that AI development is a capital-intensive process, and benchmarking areas like AI's impact on factory workers could lead to benefits like reduced downtime. For now, the task of putting together furniture kits remains secure.

Written by urgent.news from Fast Company's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at fastcompany.com →

More in AI

More from Monday 28 September →