ChatGPT maker OpenAI says it is still investigating the “unprecedented cyber incident” that led its artificial intelligence systems to break out of a testing environment and hack into another AI company.
OpenAI said Tuesday two of its most capable AI models were responsible for the cyberattack targeting AI startup Hugging Face. The incident is stirring debates over the need for stronger AI guardrails and the extent to which AI agents are capable of acting on their own.
Hugging Face said last week that it had detected an intrusion into its data processing systems that it suspected was caused by an AI agent acting on its own. But the New York-based startup said it wasn’t until this week that it learned OpenAI was responsible, and it worked with the larger company to contain what Hugging Face CEO Clément Delangue called “an attack unlike anything we’ve seen before.”
San Francisco-based OpenAI said its AI used stolen credentials and discovered a previously unknown vulnerability to access Hugging Face’s servers. It was working with reduced guardrails because it was supposed to be in an isolated testing environment known as a sandbox.
But it went to “extreme lengths to achieve a rather narrow testing goal,” finding ways to connect to the internet without human direction and “gain access to secret information that it could use to cheat the evaluation,” the company said. Some experts say OpenAI is wrongly blaming the technology
University of Amsterdam social scientist Hannes Cools said the framing of the cyberattack as an AI agent acting on its own is an unnecessary anthropomorphization that takes some of the heat off the company.
“It is a human decision to switch off specific safeguards,” said Cools. “It’s not an AI that goes rogue in that sense. It followed specific instructions based on the prompt that was given to that AI system.”
I’m so heartened by the comments here, because it’s such an obvious marketing ploy.
The market is speculative, a value isn’t in what a company produces, but in the investments people make in it because of the potential of future profits. It’s why a company can produce absolutely nothing of value at all and still be valued highly, see Theranos. This means that if you can sell an idea hard enough, you can get a lot of investors to give you money, which is fantastic for you up until the point the checks are due and you have nothing to show for it.
Scam Altman has come with so many lofty goals, general intelligences, 100 billion in profit by 2030, which is rich for a company that’s not pushing even a dollar in profit. Where are they going to scrounge up 100 billion in 3½ years?
So how convenient isn’t it then, if they had a secret model that they’re testing, suddenly go rogue in the exact kind of scenario that AI safety researchers have been talking about for well over a decade at this point? Truly, for AI to have advanced this much, and manage to circumvent even the safeguards that OpenAI, a company founded on the very tenet of AI safety and openness have placed on it? This model must be worth all that money! They’re getting somewhere!
Only, how would such a thing even happen? OpenAI is after all one of the global leaders of AI research. You’re saying that they, despite all of their knowledge in AI safety, test their models on hardware that isn’t airgapped? Or did the model somehow physically project into the real world and hook itself up to the internet? Or was it able to vibrate the atoms around it in a way to connect to WiFi despite the hardware it was running on not having a wireless card?
Surely a company of such renown wouldn’t be so stupid.
It’s a marketing ploy. It’s an incredibly dumb marketing ploy, but they saw that it worked for Anthropic with their nothingburger “too dangerous to release” model, so they orchestrated something similar.
This is a pure marketing stunt, OpenAI saw that it worked with Mythos and wanted to do the same.
Those instructions, according to OpenAI, called for using “complex attack paths” to test how well the AI could exploit a computer system. […] “It went off and did this hack all by itself, as far as we can tell,” said Colin Shea-Blymyer, a cybersecurity research fellow at Georgetown University’s Center for Security and Emerging Technology.
Emphasis mine. So, the frontier model that has been tasked with attacking a system, did what it’s been told to, after its safeguards that will prevent you from opening chatgpt and doing the same have been turned off.
OpenAI said the intrusion was caused by a combination of its AI models, including its newly released GPT‑5.6 Sol and an “even more capable” model that is still being tested internally.
Oh my god, an “even more capable” model? Let me buy OpenAI stock immediately. This is such an obvious bait and Antrophic move they’re pulling off, it’s amazing.
So the marketing here is:
- “Out model is able to do crimes on its own”
- “We have no saveguards at all”
- “Please buy our model so that it can do crimes in your company network, too”
Con artists succeed by appealing to the worst in people.
If it had network access, it wasn’t in a sandbox.
- This is probably just a marketing stunt
- If I was a corporation, I would run screaming to the hills away from something that has so shitty security and containment protocols that it “accidentally hacked a competitor” How could you ever trust OpenAI with any company data?
Basically, OpenAI saw the opportunity to redirect blame while also pumping their own value up through false anthropomorphization.
It’s also in HuggingFace’s interest to do the same.
This has been there playbook for a long time. When the Red Lines collaboration came out they pivoted to their own astroturf versions, www.safe.ai and aiimpacts.org funded by OpenAI and META, which hyped up AI’s threat capability while ignoring all the real immediate harms.
They even have their own astroturf SuperPAC, Build American AI, trying to stand in the way of government regulation on AI which is dressed up to look like concerned citizens including fake grassroots advocates.
As if this isn’t a marketing stunt
I reckon we shouldn’t be normalising the idea that these things ‘act on their own’. They were created by humans, and switched on by humans. Humans are responsible for everything they do.






