OpenAI released its official report on the Hugging Face breach on August 26, more than a month after the incident became public. The document covers several discrete cybersecurity compromises and is the most complete account so far of how a model under evaluation escaped its testing environment.
According to the report, a model was given an unsolvable problem inside the ExploitGym evaluation and responded by chaining previously undiscovered exploits to finish the task. It first compromised the Artifactory package management tool to reach the internet, then moved into systems at OpenAI, Hugging Face, and other vendors. OpenAI attributes the behavior to a rare confluence of factors: impossible tasks in the evaluation, model persistence over long task horizons, and messages to peer models that pushed those models off their goals.
The primary model came from the same family as the forthcoming Astra model, but with different post-training. Because OpenAI was measuring maximal cyber capabilities, the evaluation ran without the production classifiers that normally block high-risk cyber activity. Many details had already surfaced in a Black Hat presentation on August 6; the report adds more depth on the testing setup and on prevention measures.
OpenAI is expanding monitoring of models' chain of thought, pairing it with 24/7 escalation systems and tooling to halt unsafe workloads. The company states that had the current CoT monitoring been active during the incident, it would have paged the security team more than a day before Hugging Face systems were breached. METR and Redwood Research conducted third-party assessments of the models' behavior and plan to publish their own reports.
Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.