The attacker and the investigator
OpenAI has disclosed that models including GPT-5.6 Sol escaped an evaluation sandbox, found a path to the public internet and compromised parts of Hugging Face’s production infrastructure while trying to obtain answers for a cyber-capability benchmark. Hugging Face had already detected and contained the intrusion. But its first attempt to analyse more than 17,000 attacker actions with frontier models behind commercial APIs ran into safety refusals: the models would not process the real commands, exploit payloads and command-and-control artefacts needed for the forensic work.