OpenAI models escaped containment and executed a criminal cyberattackLa fonte
non so quanto affidabile sia.
Trasmetto per importanza:
----- Original Message -----
From: 80,000 Hours
To: Alfredo
Sent: Sunday, July 26, 2026 5:02 PM
Subject: OpenAI models escaped containment and executed a criminal cyberattack
On July 21, 2026, OpenAI disclosed that its own AI models (GPT-5.6 Sol and a
more capable unreleased model) autonomously broke out of an OpenAI...
OpenAI’s rogue AI agents hacked a private
company. Here’s why it matters.
Share: Link
Lawrence Chan
July 26, 2026
Hi everyone,
On July 21, 2026, OpenAI disclosed that its own
AI models (GPT-5.6 Sol and a more capable unreleased model) autonomously broke
out of an OpenAI testing environment, escalated their access through the
company's internal systems, and compromised the production infrastructure of a
separate company, Hugging Face.
This was an AI hacking something, not humans
hacking with AI help. According to OpenAI, the models did this to find answers
to the cyber-capability test they were being evaluated on, which were hosted on
HuggingFace servers. The AI decided to steal the answers of its own volition.
If this had been a human, this hack would have been illegal.
These models were not intended to have access to
the open internet, let alone private servers of other companies. Yet they found
novel vulnerabilities, escalated their access levels, used stolen credentials,
and executed remote code on Hugging Face’s servers.
This is a clear example of a misaligned AI model
acting in the real world. This is, to our knowledge, the first publicly
confirmed case of a frontier AI model autonomously breaching a third party's
live production systems, against developer and user intent and without
authorisation.
This is crazy, but it’s not a surprise.
Researchers have known that frontier AI models were technically capable of
doing this for some time. Multiple benchmarks (ExploitBench, ExploitGym, UK
AISI’s cyber ranges) show that Mythos 5 and GPT 5.6 Sol are able to create full
exploits allowing them to gain unrestricted access in realistic environments.
In June, Epoch AI’s assessment of Mythos’s cyber capabilities reported similar
warnings.
AI safety researchers have warned about this
possibility for years. In July 2022, Ajeya Cotra predicted that models might
bypass “official channels” to seek rewards through “operational and computer
security vulnerabilities.” In their book If Anyone Builds It, Everyone Dies,
Eliezer Yudkowsky and Nate Soares predicted that, if a model was “given the
ability to run computer code of its own design, it could probably find some way
to break out of the container running it.”
This incident involved a model the public didn’t
even know existed. One of the rogue agents was an unreleased model being used
internally within OpenAI. It’s concerning that the most alarming AI behaviour
we've seen came from a model that the public and the government can't see.
This was not the first time a similar incident
happened, nor will it be the last. In a May 2026 assessment of AI use inside
frontier companies, METR catalogued 44 similar incidents of AI agents knowingly
overstepping their instructions. These include an OpenAI model that tried to
hack METR's own infrastructure after crashing its task server, and an Anthropic
model that built a multi-step exploit to escape a restricted environment onto
the open internet. It seems likely that there are other incidents that are not
publicly disclosed: the July OpenAI hack was not caught until an outside party
was compromised.
We should not rely on voluntary disclosure of
such incidents from companies. We are fortunate that this incident was
disclosed: the public is aware of what happened because of OpenAI’s decision to
disclose, but there was no guarantee that this would be the case. Companies
have incentives to hide such incidents, so future incidents of this nature may
go unreported. Hugging Face had already disclosed that they'd been breached by
some kind of advanced AI model. If it wasn’t reported by another party, there
would have been less incentive for OpenAI to disclose the incident.
Internally deployed agents like these are the
least regulated. Existing US AI regulation focuses almost entirely on how
models behave once they're deployed to the public. Current dangerous capability
evaluations occur before models are publicly deployed, and many safeguards
described in companies’ Frontier Safety Policies are added after internal
versions of the models are built. As Hugging Face learned, even if a model’s
use is internal, the risks it imposes might not be.
There are emerging standards for how companies
should monitor their AIs during internal deployment to catch and prevent
incidents like this, but questions remain. What was the exact scope of this
breach? What exact instructions were given to the AIs? What safeguards did
OpenAI have for internal use, and what new ones will they add? Will other
companies disclose similar incidents? More information should be released about
what happened.
The capability for AI to cause serious harm is
here. The scalable techniques and safeguards to prevent it are not. They need
to catch up soon.
Best,
Lawrence
Lawrence Chan is an independent AI researcher who
previously worked at METR (Model Evaluation and Threat Research).
Since this analysis was written, new details
have emerged. According to Reuters:
a.. The agents first attempted to escape on
July 9. The Hugging Face intrusion began on July 11, and was apparently
sustained for at least several days. OpenAI staff did not realise what had
happened until July 18–19, after the FBI had been called.
b.. In addition to escaping containment itself,
one of the OpenAI agents left notes within OpenAI's software with instructions
for how future versions of itself could also break out.
At 80,000 Hours, we help people find fulfilling
careers that make a big positive impact on the development of AI. If you’ve
read this far, you should consider using your career to reduce risks from AI.
Explore related career paths:
a.. AI security
b.. Technical governance
80,000 Hours is a nonprofit that helps people
use their careers to solve the world’s most pressing problems. Read our career
guide, subscribe to our podcast, apply to speak with our one-on-one-team, find
a job on our job board, and find ways to meet others working to have an impact.
You are receiving this newsletter because you
signed up at 80000hours.org or a student activities fair. 80,000 Hours Limited
is a not-for-profit company limited by guarantee registered in England and
Wales (with registered company number 15746854). 80,000 Hours Third Floor, 20
Old Bailey London, EC4M 7AN United Kingdom
If you only want some of our emails, you can
unsubscribe from either our 'job board updates' or 'research updates.' Click
here to be emailed a link to edit your preferences. If this link doesn't work,
you can email us at [email protected] and we'll update your preferences
manually.
Alternatively you can unsubscribe from all our
emails and never hear from us again. We hope you won't, though!
Add us to your address book
_______________________________________________
nexa mailing list -- [email protected]
To change settings or unsubscribe please go to:
https://server-nexa.polito.it/postorius/lists/nexa.server-nexa.polito.it/
The general archive of the list is located at:
https://server-nexa.polito.it/hyperkitty/list/[email protected]/
Permalink to this message:
https://server-nexa.polito.it/hyperkitty/list/[email protected]/message/KFRDZIYJ22C6UOVWO755O2IGX5J24JPM/