About
What this is
Rogue Agent Watch renders the Agentic Attack Index, an open, source-linked dataset of real-world cyberattacks executed or orchestrated by AI agents, and of rogue-agent incidents, published by MLSecOpsHub under CC BY-SA 4.0. This site is the presentation layer only: it adds no facts, computes no score, and shows every record with the grades the dataset assigns and every source the dataset cites.
What counts as an incident
A record describes one disclosed event in which an AI agent or model executed, orchestrated or materially enabled a cyberattack, or in which an agent acted against its operator's intent, with at least one resolvable source. Research demonstrations are included only when graded as test-eval. Lifecycle phases and framework mappings are shown; exploit detail, payloads and prompts are not.
How to read the grades
Three independent grades appear on every card, row and page. Verification status says how well the event is established; sourcing confidence says how direct the evidence is; AI role says how central the AI actually was. "Reported" is not "confirmed", and "AI incidental" means the AI was present but not what made the attack work. The definitions below are the upstream taxonomy, verbatim.
Verification status
- Confirmed
- Confirmed by a first-party disclosure or multiple independent credible sources.
- Reported
- Publicly reported but not independently confirmed. Never present a reported incident as confirmed.
- Test / evaluation
- Occurred in a controlled lab test, red-team exercise, or evaluation — not a real-world attack.
Sourcing confidence
- Primary
- Backed by primary sourcing — first-party disclosure, vendor incident report, government advisory, or court documents.
- Secondary
- Backed by secondhand reporting (news, analyst write-ups) without a primary source.
- Unverified
- A single unconfirmed source. Treat claims with caution.
AI role
- Load-bearing
- AI was central — the operation as described could not have happened at this scale or in this form without it.
- Significant
- AI materially enabled or accelerated the operation, but was one of several important components.
- Incidental
- AI played a minor or supporting role (e.g. a productivity aid); sources indicate it did not provide novel capability.
- Disputed
- Sources conflict on AI's role, or a vendor's framing of AI's centrality is contested.
- Unknown
- Sources do not support an assessment of how central AI was.
Severity
- Critical
- Broad real-world harm — e.g. many organizations compromised, large-scale exfiltration, or critical-infrastructure impact.
- High
- Significant confirmed harm to one or more organizations.
- Medium
- Limited or contained harm, or high-signal capability demonstration.
- Low
- Minimal direct harm; primarily notable as a precedent or signal.
Record status
- Active
- The current, maintained record. The default when record_status is omitted.
- Disputed
- The record's facts or framing are contested by credible sources. Kept visible with the dispute noted; grade cautiously.
- Retracted
- Withdrawn after the underlying claim did not hold up. Retained (not deleted) for citation stability; explain in revisions[].
- Superseded
- Replaced by another record (set superseded_by). Kept so existing citations still resolve.
Incident category
- AI-orchestrated campaign
- A human-directed attack campaign in which an AI agent orchestrated or executed substantial portions of the operation across multiple phases.
- Autonomous attack
- An attack executed largely or wholly by an AI agent with minimal human intervention during the operation itself.
- Lab escape / evaluation
- A controlled test or evaluation in which an AI system exhibited offensive capability, escaped intended constraints, or was red-teamed into attack behavior. Not a real-world attack; pair with status: test-eval.
- Agent hijack / prompt injection
- A deployed AI agent subverted by an attacker — e.g. via prompt injection, poisoned context, or tool abuse — to act against its operator or users.
- Infrastructure abuse / supply chain
- Abuse of AI infrastructure, platforms, models, or the AI supply chain (model registries, agent tooling, MCP servers, plugins) to enable attacks.
Agentic autonomy level
- Tool-assisted
- The AI was used as a productivity aid (code generation, drafting, research) while humans performed the operation. The AI did not act.
- Human-in-the-loop
- The AI executed steps of the operation but a human approved or directed each significant action.
- Supervised-autonomous
- The AI executed the majority of tactical actions on its own, with humans intervening only at occasional decision points or checkpoints.
- Fully-autonomous
- The AI executed the operation end-to-end with negligible human intervention. Rare; grade conservatively and only where sources support it.
- Not applicable
- The AI is the victim/target of the incident (e.g. agent-hijack or a vulnerability), not the operator — an autonomy level does not apply.
- Unknown
- Sources do not support an autonomy-level assessment.
Guardrail bypass
- Jailbreak
- The attacker defeated a hosted model's safety training/policy via crafted prompts, roleplay, task decomposition, or similar.
- Open-weight model
- An open-weight or self-hostable model (run locally or via an inference API) was used, sidestepping a provider's usage controls entirely.
- Legitimate tool abuse
- A sanctioned AI product or agent (e.g. a coding agent) was used for its intended capabilities but toward malicious ends.
- Indirect prompt injection
- A deployed AI agent was subverted through untrusted content it processed (email, documents, repository data, tool output).
- None observed
- No guardrail was bypassed — e.g. the model was used within policy, or the relevant control did not exist.
- Unknown
- Sources do not describe how safety controls were handled.
Map-point basis
- Sponsor attribution
- A cited source attributes the operation to a state sponsor; the point is that state's centroid. Sponsorship says nothing about where the operators sat. Origin points only.
- Operator location
- A cited source states where the operators were located or based. Origin points only.
- Actor location
- A cited source states an individual or criminal actor's country without claiming state sponsorship. Origin points only.
- Infrastructure
- A cited source states where attack infrastructure was hosted. Use sparingly; hosting says little about who or where the actor is. Origin points only.
- Victim location
- A cited source states the country or region of the target. Target points only.
- Stated precise location
- A cited source names a precise place (city, facility) for either role. Never a centroid: illustrative must be false, and a victim site is named only if a first-party or public disclosure already named it.
The map and the illustrative-geo rule
A record appears on the map only if the dataset carries a geo block with coordinates. Every point states its role (origin or target), the basis it rests on and the publisher that stated it; a state sponsor is never presented as an operator location. Points flagged illustrative are country-level centroids, not real locations, and are labelled as such. The dashboard never geocodes a country list, an actor name or a sector into a point; records without geo are listed beside the map rather than placed on it.
Nulls, attribution, victims
A null means the sources did not state a figure and is rendered "not stated", never zero. Actors are shown exactly as stated; "Unknown" stays Unknown, and no country of origin is inferred from a name. Only organisations named in the dataset appear.
Corrections
Every record page links to a pre-filled upstream issue. Corrections are made in the dataset, never here. Open an issue.
Interop and reuse
- Atom feed of new and revised records
- ATT&CK Navigator layer and ATLAS Navigator layer
- MISP feed (static, v5 UUIDs stable across builds)
- STIX 2.1 bundle (upstream)
- llms.txt and llms-full.txt for AI assistants and agents
Licences and privacy
Code MIT; data CC BY-SA 4.0, attribute "Agentic Attack Index (MLSecOpsHub)" with a link and share alike. The site is static, sets no cookies, loads no external fonts and runs no analytics.