Artificial intelligence systems can now detect assembly mistakes in IKEA furniture. Epoch AI created the Furniture Assembly Benchmark, known as FAB. This benchmark uses photographs of three IKEA pieces captured during assembly with intentional errors.
Tested models must compare these photos against official instructions. They must identify the mistakes and describe what went wrong. The testing measures the visual reasoning and spatial awareness of advanced models.
Progress in this benchmark has accelerated quickly. In November 2025, Claude Opus 4.5 held the top score at 28 percent. Ten months later, OpenAI's GPT-6 Astra reached an 80 percent accuracy rate.
Read nextOpenAI Pauses Training of Its Most Powerful Models After Sandbox IncidentGPT-6 Astra completes the task at a speed of three minutes per photograph. Other models also scored high marks in the evaluation. Claude Fable 5.1 achieved 70 percent, while Claude Opus 5 reached 61 percent.
Chinese open-weight models lag behind the frontrunners. Systems like Kimi K3 trail the leading AI models by at least seven months in this specific visual task. Astra also demonstrates strong performance in visual robotic tasks.
Researchers note that the technology remains too slow for real-time assembly assistance. Despite this speed limitation, the system could eventually help with car repairs or appliance fixes.
The rapid improvement shown by Astra is significant. Models struggled with much simpler visual tasks only a short time ago.
Researchers will continue evaluating AI capabilities on complex visual benchmarks as models improve.



