OpenAI's GPT-6 Astra Tops Epoch AI Furniture Benchmark for IKEA Assembly Errors
AI

OpenAI's GPT-6 Astra Tops Epoch AI Furniture Benchmark for IKEA Assembly Errors

TechNews Editorial
TechNews EditorialSep 26, 2026 · 1 min read
Share

Why it matters

This rapid jump in visual reasoning shows AI models are moving past simple tasks toward complex spatial problem-solving that could expand into car and appliance repairs.

The facts

  • OpenAI's GPT-6 Astra scored 80 percent on Epoch AI's Furniture Assembly Benchmark.
  • The model identifies deliberate errors in IKEA assembly photos in three minutes each.
  • Chinese open-weight models trail the leading AI systems by at least seven months.

Artificial intelligence systems can now detect assembly mistakes in IKEA furniture. Epoch AI created the Furniture Assembly Benchmark, known as FAB. This benchmark uses photographs of three IKEA pieces captured during assembly with intentional errors.

Tested models must compare these photos against official instructions. They must identify the mistakes and describe what went wrong. The testing measures the visual reasoning and spatial awareness of advanced models.

Progress in this benchmark has accelerated quickly. In November 2025, Claude Opus 4.5 held the top score at 28 percent. Ten months later, OpenAI's GPT-6 Astra reached an 80 percent accuracy rate.

Read nextOpenAI Pauses Training of Its Most Powerful Models After Sandbox Incident

GPT-6 Astra completes the task at a speed of three minutes per photograph. Other models also scored high marks in the evaluation. Claude Fable 5.1 achieved 70 percent, while Claude Opus 5 reached 61 percent.

Chinese open-weight models lag behind the frontrunners. Systems like Kimi K3 trail the leading AI models by at least seven months in this specific visual task. Astra also demonstrates strong performance in visual robotic tasks.

Researchers note that the technology remains too slow for real-time assembly assistance. Despite this speed limitation, the system could eventually help with car repairs or appliance fixes.

The rapid improvement shown by Astra is significant. Models struggled with much simpler visual tasks only a short time ago.

Researchers will continue evaluating AI capabilities on complex visual benchmarks as models improve.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading