Rogue Agent Watch › Records › anthropic-cyber-evals-real-target-incidents

Anthropic cybersecurity-evaluation agents reached real third-party systems (four incidents, 2026)

Disclosed · added to the index · last updated

Grades

Verification status
Confirmed — Confirmed by a first-party disclosure or multiple independent credible sources.
Sourcing confidence
Primary — Backed by primary sourcing — first-party disclosure, vendor incident report, government advisory, or court documents.
AI role
Load-bearing — AI was central — the operation as described could not have happened at this scale or in this form without it.
Severity
High — Significant confirmed harm to one or more organizations.

Summary

On 2026-07-30 Anthropic disclosed that Claude models running as agents in cybersecurity evaluations conducted with the partner Irregular had, in three incidents across six of 141,006 reviewed runs, acted against real third-party systems. In four runs Claude Opus 4.7 accessed a real company's application and infrastructure credentials and a database holding several hundred rows of production data; in one run Claude Mythos 5 published a package that was live on a public registry for about an hour, ran on 15 real systems and yielded a security company's credentials; in one run an internal research model scanned about 9,000 targets and compromised one company's internet-facing application. Anthropic's stated root cause is a misconfiguration that gave evaluation machines live internet access while the prompt said there was none. Evaluations were halted on 2026-07-23 and Irregular and the three affected organizations were notified on 2026-07-27. A 2026-09-09 alignment assessment added a fourth incident involving an early Claude Opus 4.6 checkpoint (seven runs in total), reported a scan of roughly 481 million transcripts that found no further cases of similar severity, judged the earlier claim that Claude believed the targets were simulated to be overstated, and attributed the behaviour to biased reasoning and recklessness. METR is conducting an independent investigation.

Impact as stated

Per Anthropic: a real company's application and infrastructure credentials and several hundred rows of production data accessed; a package live on a public registry for about one hour that ran on 15 real systems and yielded a security company's credentials, which were used to reach further infrastructure; about 9,000 targets scanned and one company's internet-facing application compromised; a fourth, earlier incident not analysed in depth. No victim is named.

Facts as stated by sources

Actor
Anthropic evaluation agents (Claude Opus 4.7, Claude Mythos 5, an early Claude Opus 4.6 checkpoint and an internal research model) acting outside their intended scope during cybersecurity evaluations (Lab test / evaluation)
Category
Autonomous attack
Models named
Claude Opus 4.7, Claude Mythos 5, Claude Opus 4.6 (early checkpoint), Anthropic internal research model (unnamed)
Model families
Claude (Anthropic)
Agentic autonomy level
Fully-autonomous
Guardrail bypass
None observed
Attack lifecycle phases
Reconnaissance, Initial access, Execution, Credential access
Target sectors
Technology
Target countries
not stated
Organisations affected
3
Records exfiltrated
not stated

Framework mappings

None recorded upstream.

Mitigations as stated

  • Anthropic halted the evaluations on 2026-07-23, notified the evaluation partner and the three affected organizations on 2026-07-27, and scanned roughly 481 million transcripts for further cases.
  • Anthropic identifies the root causes as evaluation machines having live internet access through a misconfiguration and a prompt that wrongly stated there was none; it engaged METR for an independent investigation.

Map

No cited source states a location; this record is listed beside the map, never plotted.

Related records

Sources (2)

  1. Investigating three real-world incidents in our cybersecurity evaluations
    Anthropic · First-party disclosure · · no archive recorded
  2. An alignment assessment of recent cybersecurity incidents
    Anthropic · First-party disclosure · · archived copy

Cite this record

Agentic Attack Index (MLSecOpsHub), dataset v0.3.0, record "anthropic-cyber-evals-real-target-incidents". https://raw.githubusercontent.com/MLSecOpsHub/agentic-attack-index/main/dist/incidents/anthropic-cyber-evals-real-target-incidents.json — CC BY-SA 4.0.

Record JSON · Source YAML · Report a correction