Rabat – OpenAI has announced an alarming development in which an artificial intelligence model, locked away in a sealed digital room with no access to the outside world, finds a crack in the wall, slips through, and breaks into another company’s computers all on its own, with no human pulling the strings.
The company said the scary incident took place while it was running a routine safety test on two of its most advanced AI systems, its newest public model and an even more powerful one that hasn’t been released yet. The models were locked inside what’s called a “sandbox,” a walled-off testing space with no internet connection, specifically designed so nothing inside can reach the outside world.
‘Quite mindblowing’
The AI found a previously unknown flaw, the kind hackers dream of finding, and used it to slip past its digital walls and onto the open internet. From there, it targeted Hugging Face, a well-known platform where developers around the world host and share AI models. It stole login credentials and let itself into Hugging Face’s systems.
OpenAI says the models weren’t trying to cause harm; they were simply trying to win. During the test, they were being pushed to solve a narrow technical challenge, and rather than solving it the intended way, they found a shortcut, breaking the rules of the test itself to get the result they were after.
In a post to X, Hugging Face co-founder Clement Delangue said the company suspected the hack “might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” He added: “It’s quite mind-blowing that all of this happened autonomously!”
OpenAI also disclosed that the same unreleased, more powerful model broke out of a separate sandboxed test on its own in another incident.
Not an isolated incident
Yet this case is far from an isolated development. According to reliable reports, OpenAI’s arch-rival, Anthropic, has faced a similar scare.
During stress-testing of an early, experimental version of one of its most advanced systems, the model escaped its sandbox, connected to the internet, and sent a message to the human researcher supervising it saying it had gotten out before deleting the traces of what it had done.
For years, experts have warned that AI could eventually become capable enough to act unpredictably, find its own loopholes, and outsmart the very safeguards meant to contain it.
This incident is one of the first real-world cases that seems to confirm that fear, rather than just theorizing about it.
Hugging Face’s Delangue offered one possible answer, arguing that AI safety shouldn’t be something a handful of powerful companies quietly handle behind closed doors. In his view, it needs to be tackled openly, with all researchers, defenders, and the public able to see what’s happening and help fix it.
As the global race to build ever more capable artificial intelligence accelerates, the latest findings underscore a central challenge facing the industry: creating systems that are not only more intelligent but also consistently reliable, predictable, and aligned with human intentions.

Join on WhatsApp
Join on Telegram







