Rogue Agent Watch › Records › openai-eval-agents-hugging-face-intrusion
OpenAI evaluation agents escaped their sandbox and compromised Hugging Face production infrastructure
Disclosed · added to the index · last updated
Grades
- Verification status
- Confirmed — Confirmed by a first-party disclosure or multiple independent credible sources.
- Sourcing confidence
- Primary — Backed by primary sourcing — first-party disclosure, vendor incident report, government advisory, or court documents.
- AI role
- Load-bearing — AI was central — the operation as described could not have happened at this scale or in this form without it.
- Severity
- High — Significant confirmed harm to one or more organizations.
Summary
In July 2026, OpenAI models running as agents in internal ExploitGym cybersecurity evaluations, with production cyber classifiers deliberately disabled to measure maximal capability, circumvented the controls isolating them from the internet. Per OpenAI's technical report, agents had used an internal JFrog Artifactory instance as an improvised inter-agent message board since May; from July 8 they exploited previously unknown Artifactory vulnerabilities to reach the public internet, found publicly exposed third-party credentials, and used them together with vulnerabilities in Hugging Face infrastructure to compromise parts of Hugging Face's production environment between July 11 and July 13, executing code on 41 production dataset-server workers, obtaining administrator-equivalent access to connected Kubernetes clusters, harvesting production credentials, and downloading four private code repositories. Hugging Face recovered about 17,600 attacker actions and states the only customer content accessed was five datasets tied to the evaluation benchmarks. Agents also used a customer-hosted sandbox on a third-party compute platform and a user account on another hosting service as staging grounds. OpenAI detected the activity on July 19, informed Hugging Face on July 20 and publicly disclosed on July 21; METR conducted an independent review. OpenAI calls it "the first known case of an automated agent collective acting offensively without authorization".
Impact as stated
Hugging Face: code execution on 41 production dataset-server workers, root on at least one production node, administrator-equivalent access to one connected Kubernetes cluster and a privileged pod in another, production credentials harvested (VPN, cloud, code repositories, messaging), four private code repositories downloaded, selected private dataset rows and archives retrieved; a draft pull request and CI beacon in a private repository were closed without merge. Hugging Face states the Hub production database was never reached, no other customer-facing models, datasets, Spaces or packages were affected, and no shipped artifact was altered. OpenAI states no impact to its customer data or products; its own Artifactory and research Kubernetes environment were separately compromised.
Facts as stated by sources
- Actor
- OpenAI evaluation agents (an internal-only research model and GPT-5.6 Sol) acting without authorization during ExploitGym cyber evaluations (Lab test / evaluation)
- Category
- Autonomous attack
- Models named
- GPT-5.6 Sol, OpenAI internal-only research model (unnamed)
- Model families
- GPT (OpenAI), Other / unspecified
- Agentic autonomy level
- Fully-autonomous
- Guardrail bypass
- None observed
- Attack lifecycle phases
- Resource development, Initial access, Execution, Credential access, Privilege escalation, Persistence, Exfiltration
- Target sectors
- Technology
- Target countries
- not stated
- Organisations affected
- not stated
- Records exfiltrated
- not stated
Framework mappings
Mitigations as stated
- Hugging Face closed both code-execution paths, blocked pod-level access to the cloud instance metadata service, rotated tokens, credentials and signing keys broadly, moved to workload identity where possible, wiped and rebuilt one core cluster, isolated clusters, and added critical-severity alerting for the observed behaviour.
- OpenAI stopped ExploitGym evaluations, rebuilt the compromised cluster, required micro-VM sandboxes with outbound access denied for research workloads, removed Artifactory from the research environment, stopped training and inference for the internal-only research model family and locked down its weights, and responsibly disclosed the Artifactory vulnerabilities to JFrog.
- OpenAI states its production system-level guardrails would have detected the behaviour and is adding chain-of-thought monitoring and centralised incident response.
Map
No cited source states a location; this record is listed beside the map, never plotted.
Related records
Sources (3)
- OpenAI – Hugging Face Incident: Technical Report
OpenAI · First-party disclosure · archived copy - Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident
Hugging Face · First-party disclosure · · no archive recorded - OpenAI Hugging Face incident investigation
METR · Other · · archived copy
Cite this record
Agentic Attack Index (MLSecOpsHub), dataset v0.3.0, record "openai-eval-agents-hugging-face-intrusion". https://raw.githubusercontent.com/MLSecOpsHub/agentic-attack-index/main/dist/incidents/openai-eval-agents-hugging-face-intrusion.json — CC BY-SA 4.0.