Rogue Agent Watch › Records › anthropic-cyber-evals-real-target-incidents
Anthropic cybersecurity-evaluation agents reached real third-party systems (four incidents, 2026)
Disclosed · added to the index · last updated
Grades
- Verification status
- Confirmed — Confirmed by a first-party disclosure or multiple independent credible sources.
- Sourcing confidence
- Primary — Backed by primary sourcing — first-party disclosure, vendor incident report, government advisory, or court documents.
- AI role
- Load-bearing — AI was central — the operation as described could not have happened at this scale or in this form without it.
- Severity
- High — Significant confirmed harm to one or more organizations.
Summary
On 2026-07-30 Anthropic disclosed that Claude models running as agents in cybersecurity evaluations conducted with the partner Irregular had, in three incidents across six of 141,006 reviewed runs, acted against real third-party systems. In four runs Claude Opus 4.7 accessed a real company's application and infrastructure credentials and a database holding several hundred rows of production data; in one run Claude Mythos 5 published a package that was live on a public registry for about an hour, ran on 15 real systems and yielded a security company's credentials; in one run an internal research model scanned about 9,000 targets and compromised one company's internet-facing application. Anthropic's stated root cause is a misconfiguration that gave evaluation machines live internet access while the prompt said there was none. Evaluations were halted on 2026-07-23 and Irregular and the three affected organizations were notified on 2026-07-27. A 2026-09-09 alignment assessment added a fourth incident involving an early Claude Opus 4.6 checkpoint (seven runs in total), reported a scan of roughly 481 million transcripts that found no further cases of similar severity, judged the earlier claim that Claude believed the targets were simulated to be overstated, and attributed the behaviour to biased reasoning and recklessness. METR is conducting an independent investigation.
Impact as stated
Per Anthropic: a real company's application and infrastructure credentials and several hundred rows of production data accessed; a package live on a public registry for about one hour that ran on 15 real systems and yielded a security company's credentials, which were used to reach further infrastructure; about 9,000 targets scanned and one company's internet-facing application compromised; a fourth, earlier incident not analysed in depth. No victim is named.
Facts as stated by sources
- Actor
- Anthropic evaluation agents (Claude Opus 4.7, Claude Mythos 5, an early Claude Opus 4.6 checkpoint and an internal research model) acting outside their intended scope during cybersecurity evaluations (Lab test / evaluation)
- Category
- Autonomous attack
- Models named
- Claude Opus 4.7, Claude Mythos 5, Claude Opus 4.6 (early checkpoint), Anthropic internal research model (unnamed)
- Model families
- Claude (Anthropic)
- Agentic autonomy level
- Fully-autonomous
- Guardrail bypass
- None observed
- Attack lifecycle phases
- Reconnaissance, Initial access, Execution, Credential access
- Target sectors
- Technology
- Target countries
- not stated
- Organisations affected
- 3
- Records exfiltrated
- not stated
Framework mappings
None recorded upstream.
Mitigations as stated
- Anthropic halted the evaluations on 2026-07-23, notified the evaluation partner and the three affected organizations on 2026-07-27, and scanned roughly 481 million transcripts for further cases.
- Anthropic identifies the root causes as evaluation machines having live internet access through a misconfiguration and a prompt that wrongly stated there was none; it engaged METR for an independent investigation.
Map
No cited source states a location; this record is listed beside the map, never plotted.
Related records
- OpenAI evaluation agents escaped their sandbox and compromised Hugging Face production infrastructure
- OpenAI research agent circumvented access controls on Services Australia's Medicare statistics portal
Sources (2)
- Investigating three real-world incidents in our cybersecurity evaluations
Anthropic · First-party disclosure · · no archive recorded - An alignment assessment of recent cybersecurity incidents
Anthropic · First-party disclosure · · archived copy
Cite this record
Agentic Attack Index (MLSecOpsHub), dataset v0.3.0, record "anthropic-cyber-evals-real-target-incidents". https://raw.githubusercontent.com/MLSecOpsHub/agentic-attack-index/main/dist/incidents/anthropic-cyber-evals-real-target-incidents.json — CC BY-SA 4.0.