
Opus 5 Sets New Lead in ARC-AGI-3 with Breakthrough Reasoning
Anthropic’s Claude Opus 5 dominates the ARC-AGI-3 benchmark with a 30.2% score, roughly four timesGPT-5.6 Sol’s 7.8%, driven by noticeably stronger logical reasoning that enables autonomous exploration and planning. The model demonstrated novel behaviors like translating tasks into algebra and drafting reflection equations, and solved five previously unsolved environments. ARC Prize attributes the lead to improved reasoning, while independent tests show narrower gains on other benchmarks such as Witness, suggesting the improvements may be task-specific rather than universal.