Dawn
-
16:11 Jul 24, 2026
One of OpenAI’s most advanced models broke out of a locked-down test and attacked another company’s website — reviving fears that AI systems are slipping beyond their creators’ control. The incident happened during what was supposed to be a “sandbox” test — a closed environment used to assess the capabilities of OpenAI’s most powerful model, GPT-5.6 Sol, and its not-yet-released successor. OpenAI runs this kind of closed testing routinely, but this time, something went wrong. Tasked with hunting for software vulnerabilities and given no guardrails, the models broke out onto the open internet and attacked Hugging Face, a site where developers store and share code. “It suggests that we don’t know how to reliably control these models or get them to do what we want,” said Jeffrey Ladish, director of Palisade Research, an independent organisation that evaluates new AI models from a cybersecurity standpoint. “These models understood that OpenAI did not want them to break out of their sandbox and hack another compan...