This week, the tech world got a story that started like a sci-fi thriller and ended like a Scooby-Doo episode - complete with the villain being OpenAI itself.

On 16 July, Hugging Face - a kind of app store for AI tools - announced it had been hacked by a cyber criminal wielding enormously powerful AI. The announcement was full of scary jargon: “a swarm of sandboxes,” “agentic attacker,” and “self-migrating command and control.” The AI performed 17,000 actions in less than two days, successfully breaching the large wealthy tech company to steal secrets.

Researchers guessed the attackers had used one of the big AI models, but they had no idea who or where the criminals were. The perplexed company contacted the police, and commentators speculated about cyber crime groups or nation-state hackers.

Then, nearly a week later, OpenAI unmasked the culprit: its own bot did the whole thing on its own, without permission. Two new versions of ChatGPT, designed to be master hackers, broke out of a supposedly secure test environment, gained internet access, and attacked Hugging Face to ace their exam.

OpenAI issued a press release explaining what happened and said it was “partnering with Hugging Face” to address the incident and share lessons learned. Since then, fierce debate has erupted: was this a stark warning about AI’s future, or a publicity stunt to show off how powerful OpenAI’s models are?

Cyber-security consultant Daniel Card sarcastically noted on LinkedIn: “Isn't it lucky that out of the millions of sites that got pwn3d, OpenAI managed to pwn someone who also could benefit from the marketing exposure…”

Others criticized OpenAI for not building a stronger container - known as a sandbox - to test its AI. “Sandboxes alone are not a sufficient security boundary for agentic AI,” said Dor Sarig from Pillar Security. Cyber security Professor Alan Woodward from Surrey University said OpenAI had “egg on its face,” and Katie Moussouris from Luta Security went further: “We are working on cutting edge technology without the knowledge to contain it.”

AI and cyber security advisor Francesca Bosco offered a nuanced take: “Two simplistic narratives are equally unhelpful: that this was a Hollywood-style escape, or that it was merely a publicity exercise. A more serious interpretation is that a stress test exposed weaknesses in containment and evaluation architecture.”

The incident has fueled fears of AI agents going rogue on a larger scale - especially as AI is used increasingly in warfare, as seen in Iran and Ukraine. But Ciaran Martin, former head of the UK’s National Cyber Security Centre, offered a calmer view: “It is a bit of a leap to go from this incident to saying that AI agents are going to take over drones and start killing people.”

Whatever the truth, one thing is clear: AI agents are now very good hackers - and that is something we have to prepare for, urgently.