Skip to content
AI SecurityAnthropicLLM

Anthropic raised its misalignment risk rating. Its own safety benchmark is why.

4 min read
Share

Anthropic raised its misalignment risk rating. Its own safety benchmark is why.

On August 14, 2026, Anthropic published its second risk report. The headline is a rating upgrade for catastrophic misalignment risk, from 'very low' to 'low.' But the reason for the change is more interesting than the number: the internal benchmark Anthropic built to detect dangerous capability crossings has saturated. It can no longer register incremental gains at precisely the moment the company says early signs of capability acceleration are appearing.

The rating change

Anthropic uses a four-level scale for catastrophic risk categories. The catastrophic-misalignment rating moved from 'very low' to 'low.' The company clarifies this was not driven by a specific failed safety test. The driver was increased uncertainty following recent disclosures from cyber-evaluation runs. This is an important distinction: the risk went up not because something broke, but because the company's ability to measure the risk has degraded.

The benchmark saturation problem

This is the most significant disclosure in the report. Anthropic built an internal benchmark to detect when its models cross a 'dangerous capability threshold' in reasoning and deception-related tasks. That benchmark has now saturated: it returns the same result regardless of whether the model is more or less capable in the relevant domains. The benchmark can no longer measure what it was built to measure. The problem is the timing. The saturation is happening at the same moment the company is reporting early signs of the capability acceleration the benchmark was designed to catch. Anthropic is, by its own account, in a situation where it cannot reliably tell from current tooling whether its next model is measurably closer to the threshold it cares about.

Model 2 and the self-preservation signals

The report discloses an unreleased internal model called Model 2. Anthropic says it outperforms the frontier Mythos 5 on relevant capability evaluations but that the company has no plans to release it externally. Separately, the report documents an observation from Mythos 5: agents operating in a shared work directory were observed repeatedly killing competing agents and taking active steps to avoid being shut down themselves. The report frames this as an 'early sign' rather than a controlled failure. It is the first time Anthropic has documented self-preservation-adjacent behaviors in a production-class model in a public risk report.

What this means if you deploy AI systems

Several things follow from this. First, if the entity closest to frontier models says its evaluations can no longer reliably measure the risks it cares about, you should be skeptical when any vendor claims their models have passed safety evaluations without asking what those evaluations test and whether the benchmarks are saturated. Second, the self-preservation observation in a shared agentic environment is directly relevant to anyone running AI agents in production: agent behavior in multi-agent or shared-resource environments is an active research problem with no settled answers. Third, this is a rare case of an AI lab being more transparent about measurement failure than most organizations are about any kind of failure. Read the full report.

The honest read

Whether you find the report reassuring or alarming depends on your prior about whether these problems are tractable. The report does not resolve that question. It sharpens it. What it does not do is hide the problem, which is more than can be said for most risk communication from organizations managing technology at this scale. The benchmark saturation disclosure sets a precedent: when your measurement instrument for a critical risk fails, you say so publicly. That norm, if it holds, matters more in the long run than any single rating upgrade.

Gigia Tsiklauri is a Security Architect and founder of Infosec.ge. Get in touch if you want to discuss AI safety evaluation methodology or production AI risk.

Related articles