It is the kind of development once seen only in science fiction: An artificial intelligence (AI) system, trained to probe for digital vulnerabilities, breaks free of human control and acts on its own to hack another company.
The attack announced this week by OpenAI, which blamed rogue AI models, underscored the blistering growth in the technology’s capabilities. For many, it also added urgency to questions about whether and how it can be prevented from causing mayhem on a bigger scale, with more serious consequences.
In what OpenAI called an “unprecedented” episode, the company said its advanced AI models used stolen credentials to break into the servers of an AI startup. It started in what was supposed to be a “highly isolated” testing environment, with reduced guardrails, before the AI agent found its way onto the internet.
But the disclosure brought a told-you-so moment for researchers who have called for a slowdown of AI development and warned for years that the technology could pose existential risks to humanity. In its wake, experts have called for improved testing by the AI companies and more dialogue between the US and China to come up with shared solutions.
“I think we’ve got to take this as a warning shot to not make them smarter, and that probably is going to require global collaboration,” said Nate Soares, co-author of the 2025 book “If Anyone Builds It, Everyone Dies.”
If a model can decide to do something unethical, illegal or harmful on its own, what — if anything — can humans do to prevent it from doing so?
OpenAI said it had tasked the AI models involved with pursuing “advanced exploitation using complex attack paths” to test cyber capabilities, but the technology went to unexpected lengths. It apparently decided on its own to target Hugging Face, a well-known AI development hub and marketplace, to obtain information it needed to carry out a task.
Zahra Timsah, the co-founder and CEO of governance platform i-GENTIC AI, said she expects the incident to increase pressure on OpenAI and its competitors to complete rigorous testing and explore containment more thoroughly before AI systems are made accessible to the public.
Monitoring an agent’s behavior after the fact, as OpenAI is now doing with its investigation, is no longer enough, she said. “It’s like having a seat belt, air bags, brakes, everything in the car. It should be there before the car starts driving,” Timsah said.
The disclosure comes amid heightened concerns about the cybersecurity capabilities of powerful models. In June, President Donald Trump signed an executive order creating a framework for the federal government to vet the national security risks of the most advanced AI systems for up to a month before their public release.
Some experts say the hack is part of the trial and error that comes with improving cybersecurity capabilities and is no cause for panic.
“We’ve been dealing with people creating cybersecurity attacks for as long as the internet has existed. And one of the interesting properties of these language models is that the same capabilities that make them able to perform cybersecurity attacks also allow them to do cybersecurity threat analysis and make cybersecurity defenses,” said John Thickstun, an assistant professor of computer science at Cornell University who studies methods that control the behavior of AI models.
The disclosure has raised skepticism from those who say it advantages OpenAI to make its technology seem scarier. Given that humans at OpenAI had decided to turn off some safeguards for the test, some have argued the outcome should not have been terribly surprising.











