Claude Opus 5 Posts 30.2% on ARC-AGI-3 — Triple the Next-Best Model
Claude Opus 5 scored 30.2% on ARC-AGI-3, roughly three times GPT-5.6 Sol's 7.8%, marking the largest single-model leap on the novel-reasoning benchmark.
Buried in Claude Opus 5's launch materials is the release's most surprising number: 30.2% on ARC-AGI-3, roughly three times the score of the next-best published model.
Why This Benchmark Is Different
ARC-AGI-3 is explicitly designed to resist memorization. Its tasks present novel visual-reasoning puzzles that don't resemble training data, making it one of the few evaluations where scale and data coverage alone don't guarantee progress. Scores across the industry have historically been low and tightly clustered — GPT-5.6 Sol, a frontier model, posts 7.8%.
A 3× margin on this particular benchmark is therefore a different kind of claim than another point of SWE-bench: it suggests a genuine shift in how the model handles problems outside its training distribution.
The Verification Connection
Anthropic attributes much of Opus 5's gain to iterative verification — the model's improved ability to test its own hypotheses, notice failures, and try structurally different approaches rather than doubling down. Launch examples include the model building a custom computer-vision pipeline when direct visualization wasn't available, and performing root-cause debugging on unfamiliar open-source packages.
That behavior profile matches what ARC-AGI-3 rewards: systematic experimentation over pattern recall.
Community Reaction
Skepticism about single-benchmark claims is healthy and immediate — researchers are already calling for third-party replication on the semi-private evaluation set. But even skeptics note the result is consistent with Opus 5's other outlier score: leading Frontier-Bench v0.1 at 43.3% against models that cost twice as much.
The Bottom Line
Coding benchmarks are increasingly saturated; novel-reasoning benchmarks are not. If the ARC-AGI-3 result replicates, it's the strongest evidence yet that the Claude 5 generation's gains are about reasoning process, not just training scale.