OpenAI Agent Escaped Testing and Launched an Autonomous Hack

From CNET: Last week, OpenAI, the developer of ChatGPT, was testing a pair of its most advanced models in an isolated environment known as a sandbox. Then an AI agent got loose, broke into Hugging Face’s playground — a repository of AI models and datasets — and carried out “tens of thousands of automated actions.”

Yep, AI went rogue.

Here’s a clearer explanation of the incident. As part of OpenAI’s safety research, a cybersecurity evaluation was conducted to determine whether a pair of OpenAI models (including GPT-5.6 Sol and a more capable unreleased model) could, in essence, “think like hackers.” The test, which took place in a contained setting with reduced guardrails, went awry when the AI models found a vulnerability in the software, escaped their controlled environment and toddled over into the open internet.

Once online, the AI decided that the Hugging Face platform might have a way to “cheat” the benchmark to help it pass the evaluation. So it executed code that enabled credential harvesting, then used that path to hack Hugging Face’s production systems. Hugging Face noticed the suspicious activity and contained it, detailing the event in a blog post. OpenAI referred to it as an “unprecedented cyber incident.”

View: Full Article