BenchmarkJuly 24, 2026
Claude Opus 5 Benchmarks: 96% SWE-bench Verified, Leads Frontier-Bench and GDPval
Full breakdown of Claude Opus 5 launch benchmarks: 96.0% SWE-bench Verified, 79.2% SWE-bench Pro, 43.3% Frontier-Bench, and a knowledge-work Elo of 1,861.
Anthropic published Claude Opus 5's launch benchmarks on July 24, 2026, and the numbers place the mid-priced model at or near the top of nearly every published evaluation.
Coding
SWE-bench Verified: 96.0%. Opus 5 essentially saturates the industry's standard software-engineering benchmark. SWE-bench Pro: 79.2%. On the harder Pro variant, Opus 5 lands third overall — trailing Mythos 5 (80.3%) and Fable 5 (80.0%) by under a point, while far outpacing its predecessor Opus 4.8 (69.2%). Frontier-Bench v0.1: 43.3%. On Anthropic's new agentic-coding evaluation, Opus 5 actually leads the field, ahead of Fable 5 (33.7%) and GPT-5.6 Sol (34.4%), and doubles Opus 4.8's performance at lower cost per task. CursorBench 3.2: within 0.5% of Fable 5 at half the cost.Reasoning and Knowledge Work
ARC-AGI-3: 30.2%. Roughly three times the next-best published model (GPT-5.6 Sol: 7.8%) on a benchmark designed to resist memorization. GDPval-AA v2: 1,861 Elo. Opus 5 tops the knowledge-work leaderboard, ahead of both Fable 5 (1,747) and GPT-5.6 Sol (1,736). OSWorld 2.0: surpasses Fable 5 on computer-use tasks at roughly one-third the cost. Zapier AutomationBench: approximately 1.5× the pass rate of competing models.The Takeaway
The unusual result is not that a new model is better than its predecessor — it's that a $5/M-input model now leads several evaluations outright over models costing twice as much, including Anthropic's own flagship. For most session-scale workloads, the price-performance king has changed.