What the benchmark measures
Epoch AI’s Furniture Assembly benchmark tests whether a model can inspect photos of three IKEA items assembled with deliberate mistakes, compare them with the instructions, and describe what went wrong.
| Model | Score |
|---|---:|
| GPT-6 Astra | 80% |
| Claude Fable 5.1 | 70% |
| Claude Opus 5 | 61% |
| Claude Opus 4.5, November 2025 | 28% |
Astra’s score is a substantial jump in ten months. The Chinese open models named in the report, including Kimi K3, trail the leaders by at least seven months.
Accuracy is not the same as assistance
Three minutes per photo makes the result hard to apply during a live furniture build. The benchmark shows that a model can find and describe an error; it does not show how quickly or reliably it can guide someone through correcting one.
The researchers point to possible uses in car and appliance repair. GPT-6 Astra also performs well on visual tasks for robots, extending the relevance of this kind of perception beyond furniture.
I think the more important measure now is not another accuracy score, but whether a model can turn visual diagnosis into timely, usable guidance. At three minutes per image, Astra has demonstrated recognition, not real-time help.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X