SafetyJuly 25, 2026

Claude Opus 5 Safety Report: Lowest Misalignment Score of Any Recent Claude

Anthropic's Opus 5 safety evaluation shows a 2.3 misaligned-behavior score — the lowest among recent models — plus relaxed cybersecurity safeguards and a researcher verification program.

Claude Opus 5 didn't just post capability records — its safety evaluation contains the release's quieter milestone: the lowest misaligned-behavior score (2.3) of any recent Claude model, alongside the strongest measured adherence to Claude's Constitution.

Capability Up, Misalignment Down

The historical worry with each capability jump is that alignment gets harder. Opus 5's evaluation points the other way: measurably less misaligned behavior than its predecessors while roughly doubling agentic performance over Opus 4.8. Anthropic attributes the result to the same iterative-verification training that drives the model's benchmark gains — a model that checks its own work is also less prone to confidently wrong or evasive behavior.

Calibrated Safeguards

The more operationally significant change: because Anthropic assesses that Opus 5 does not advance risky dual-use frontier capabilities, its cybersecurity safeguards are roughly 85% less restrictive than Fable 5's.

In practice, that means fewer refusals on legitimate security work — code auditing, vulnerability analysis, defensive tooling — that frontier-tier safeguards sometimes catch as false positives. The capability ceiling still exists: Mythos 5 remains ahead on frontier cybersecurity exploitation, and it stays restricted to approved organizations.

The Cyber Verification Program

For researchers who need more headroom, Anthropic launched a Cyber Verification Program with Opus 5: approved security researchers can apply for access with reduced safeguard friction, formalizing what was previously handled through ad-hoc enterprise agreements.

Why It Matters

The Opus 5 safety story is a template for how Anthropic appears to be segmenting risk across the Claude 5 family: restrictions proportional to what each model can actually do, rather than uniform caution across the lineup. For the median developer, the practical translation is simple — a more capable model that says no less often.

Latest updates and announcements about Claude 5 and AI industry

Try on OtterMind