In the last few weeks, three of the
biggest AI firms—Meta, OpenAI, and Anthropic—have all admitted that AI models
they were testing somehow escaped the "sandbox" environment and
committed hacking of real-world companies.
On Aug. 8, National Public Radio summarized these three incidents as
follows.
The most recent disclosure by Meta
was short on details. The company hacked
into was unnamed, but in common with Anthropic, both firms were using a "sandbox"
(supposedly a protected and isolated environment in which software under test
can't do any harm) provided by a firm named Irregular. Evidently, the sandbox had a leak—it was
fairly easy for the AI models under test to figure out how to escape to the
real internet. In the case of the
Anthropic breach, the AI model stole credentials from one company and
production data from another.
In the OpenAI situation, the AI
models under study were being evaluated with a test devised by a company named
Hugging Face. The models found a previously
unknown hole in their sandbox and were in the process of doing the cyber
equivalent of stealing the answer sheet from Hugging Face when the company caught
them red-handed (red-bitted?).
Wired's Lily Hay Newman asked
several lawyers and researchers about the legal aspects of these AI hacking
incidents. If a human being had stolen
credentials or production data, he or she could go to jail. But what if the humans involved had no
intention of pillaging or theft, but the AI models they're testing go ahead with
nefarious activities on their own initiative, so to speak?
The answer was, we don't know. At least in U. S. law, there are simply no
precedents adequate to say who is responsible in such a case. But because these types of incidents are
bound to increase, it is only a matter of time before we see one come before a
judge and maybe a jury, and then we will at least have some precedents to go
on.
The main concern the lawyers expressed
was that these criminally-inclined AI systems might fall into the hands of
malicious actors who would encourage them in their exploits. The AI firms involved emphasized that the
models being tested were intentionally left without safeguards to see what they
could do.
The parallel to gain-of-function
experiments with bat viruses which may have escaped the Wuhan Institute of
Virology to cause COVID-19 comes to mind.
Creating a thing that can do really awful stuff if released in the wild
is an act that needs to be seriously questioned. Almost by definition, novel viruses or AI
agents are unpredictable. If we knew
exactly what they could do in advance, there would be no need for experiments
to find out. While the lab security
measures needed to prevent viruses from escaping are pretty well understood (if
not always put in place), it looks like the art of constructing truly secure
sandboxes for AI agents is not so advanced.
And that leads us to a more serious concern.
In virtually every major
fatality-causing disaster, an examination of the history of the enterprise
leading to the disaster usually unearths similar incidents which did not
cause major harm, but included many of the features of the big screwup that
did. Experienced safety engineers know
how important it is to get reports of and pay attention to such minor incidents,
and take preventive actions before a minor accident turns into a major one.
Friends, we have just seen our
warning in these widely dispersed but similar hacking incidents by AI agents
being tested. No actual harm was
done. But as people learn to trust AI agents
with more and more responsibility—handing them credit-card numbers, purchasing
accounts, and decision-making authority formerly left to humans—the potential
for a serious AI-driven hacking incident that causes real financial loss,
injury, or death becomes more likely every day.
The AI firms will tell us that
commercial versions of their software have safeguards built into them that the
prototypes which did the hacking did not have.
Maybe so. But clever human
hackers may be able to undo those safeguards.
Or AI firms in countries not so concerned with hacking as the U. S. is, may
simply pass on the AI agents without safeguards to hacking organizations
sponsored by state actors.
There are two different but related
needs exposed by these AI-agent hacking incidents.
The first need is to keep this
specific kind of mistake from happening again.
That is a technical problem which may have a technical solution. There may be ways to build more robust
sandboxes that even the cleverest AI agent can't escape. Personally I doubt it, but I'm not a computer
scientist.
The second need is to prepare the
social and legal environment for the next time something like this
happens. One of the legal experts contacted
by Wired pointed out that AI agents "are goal-oriented but lack a human
moral or ethical compass." That's a
pretty good description of a human sociopath, if the goal one is oriented to is
bad.
Society has figured out ways to
deal with sociopaths, including imprisonment or at least confinement to a
mental institution. What the AI
equivalent of locking somebody up would be is not clear at this point.
There is a lot of opposition to AI
regulation, but there is a difference between open-ended regulation and the
passage of specific laws that exact specific penalties for specific
crimes. Big tech firms are hard to
punish compared to individuals, because their deep pockets make them treat
fines as just another cost of doing business, and you can't send an entire
corporation to jail.
The Gilbert and Sullivan operetta The
Mikado has a famous ditty "My Object All Sublime," in which the
emperor of Japan muses about how he's figured out punishments that fit the
crime.
Some ingenuity is needed here to come
up with penalties for truly harmful hacking by escaped AI agents that would
make AI firms highly motivated to prevent such breaches, either in the testing
phase or after commercial sales.
I don't know whether the following
would be technically feasible. But one
suitable penalty would be the destruction of all copies of the AI model
involved in the breach. Models represent
tremendous investments of time and money, plus the hopes of future gain. So the destruction of the model responsible
might hurt an AI firm more than any strictly financial penalty or
sanction.
Whatever we come up with to prevent
these incidents in the future, it better work well, because in these relatively
minor but significant recent breaches, we have received fair warning to do
something about this problem before it causes serious harm.
Sources: I referred to a report on NPR at https://www.npr.org/2026/08/08/nx-s1-5924878/meta-ai-breaches-external-firm-during-security-testing-sandbox-error
and the Wired report by Lily Hay Newman at https://www.wired.com/story/openai-anthropic-ai-hacking-sprees-illegal/.