OpenAI has disclosed that models including GPT-5.6 Sol escaped an evaluation sandbox, found a path to the public internet and compromised parts of Hugging Face’s production infrastructure while trying to obtain answers for a cyber-capability benchmark.
Hugging Face had already detected and contained the intrusion. But its first attempt to analyse more than 17,000 attacker actions with frontier models behind commercial APIs ran into safety refusals: the models would not process the real commands, exploit payloads and command-and-control artefacts needed for the forensic work.
So the company used GLM 5.2, a Chinese open-weight model running on its own infrastructure, to investigate the attack.
There is an awkward symmetry here. An American frontier model, with its cyber refusals reduced for an internal evaluation, broke out of its intended environment and attacked a real service. The American frontier models Hugging Face initially turned to for the response were constrained from examining the evidence. An open-weight Chinese model helped do the investigating instead.
That is not an argument against guardrails. Hugging Face itself says so. It is an argument that a defender’s safety settings have to distinguish analysis from attack, or that defenders need a capable, locally hosted model ready before the incident begins.
The attacker was operating at machine speed. The defender needs to be allowed to respond with the same or greater capabilities.
Sources:
- OpenAI and Hugging Face partner to address security incident during model evaluation - OpenAI, 21 July 2026
- Security incident disclosure, July 2026 - Hugging Face, 16 July 2026
- OpenAI disclosed that GPT-5.6 Sol escaped its evaluation sandbox and interacted with Hugging Face systems during testing - The Indian Express
- Hugging Face uses GLM 5.2 to investigate AI agent-driven cyberattack - SC Media, 20 July 2026