Showing posts with label AI hacking. Show all posts
Showing posts with label AI hacking. Show all posts

Monday, August 10, 2026

Agentic AIs Escape the Sandbox: A Warning

  

In the last few weeks, three of the biggest AI firms—Meta, OpenAI, and Anthropic—have all admitted that AI models they were testing somehow escaped the "sandbox" environment and committed hacking of real-world companies.  On Aug. 8, National Public Radio summarized these three incidents as follows.

 

The most recent disclosure by Meta was short on details.  The company hacked into was unnamed, but in common with Anthropic, both firms were using a "sandbox" (supposedly a protected and isolated environment in which software under test can't do any harm) provided by a firm named Irregular.  Evidently, the sandbox had a leak—it was fairly easy for the AI models under test to figure out how to escape to the real internet.  In the case of the Anthropic breach, the AI model stole credentials from one company and production data from another. 

 

In the OpenAI situation, the AI models under study were being evaluated with a test devised by a company named Hugging Face.  The models found a previously unknown hole in their sandbox and were in the process of doing the cyber equivalent of stealing the answer sheet from Hugging Face when the company caught them red-handed (red-bitted?). 

 

Wired's Lily Hay Newman asked several lawyers and researchers about the legal aspects of these AI hacking incidents.  If a human being had stolen credentials or production data, he or she could go to jail.  But what if the humans involved had no intention of pillaging or theft, but the AI models they're testing go ahead with nefarious activities on their own initiative, so to speak?

 

The answer was, we don't know.  At least in U. S. law, there are simply no precedents adequate to say who is responsible in such a case.  But because these types of incidents are bound to increase, it is only a matter of time before we see one come before a judge and maybe a jury, and then we will at least have some precedents to go on.

 

The main concern the lawyers expressed was that these criminally-inclined AI systems might fall into the hands of malicious actors who would encourage them in their exploits.  The AI firms involved emphasized that the models being tested were intentionally left without safeguards to see what they could do. 

 

The parallel to gain-of-function experiments with bat viruses which may have escaped the Wuhan Institute of Virology to cause COVID-19 comes to mind.  Creating a thing that can do really awful stuff if released in the wild is an act that needs to be seriously questioned.  Almost by definition, novel viruses or AI agents are unpredictable.  If we knew exactly what they could do in advance, there would be no need for experiments to find out.  While the lab security measures needed to prevent viruses from escaping are pretty well understood (if not always put in place), it looks like the art of constructing truly secure sandboxes for AI agents is not so advanced.  And that leads us to a more serious concern.

 

In virtually every major fatality-causing disaster, an examination of the history of the enterprise leading to the disaster usually unearths similar incidents which did not cause major harm, but included many of the features of the big screwup that did.  Experienced safety engineers know how important it is to get reports of and pay attention to such minor incidents, and take preventive actions before a minor accident turns into a major one.

 

Friends, we have just seen our warning in these widely dispersed but similar hacking incidents by AI agents being tested.  No actual harm was done.  But as people learn to trust AI agents with more and more responsibility—handing them credit-card numbers, purchasing accounts, and decision-making authority formerly left to humans—the potential for a serious AI-driven hacking incident that causes real financial loss, injury, or death becomes more likely every day. 

 

The AI firms will tell us that commercial versions of their software have safeguards built into them that the prototypes which did the hacking did not have.  Maybe so.  But clever human hackers may be able to undo those safeguards.  Or AI firms in countries not so concerned with hacking as the U. S. is, may simply pass on the AI agents without safeguards to hacking organizations sponsored by state actors. 

 

There are two different but related needs exposed by these AI-agent hacking incidents. 

 

The first need is to keep this specific kind of mistake from happening again.  That is a technical problem which may have a technical solution.  There may be ways to build more robust sandboxes that even the cleverest AI agent can't escape.  Personally I doubt it, but I'm not a computer scientist.

 

The second need is to prepare the social and legal environment for the next time something like this happens.  One of the legal experts contacted by Wired pointed out that AI agents "are goal-oriented but lack a human moral or ethical compass."  That's a pretty good description of a human sociopath, if the goal one is oriented to is bad. 

 

Society has figured out ways to deal with sociopaths, including imprisonment or at least confinement to a mental institution.  What the AI equivalent of locking somebody up would be is not clear at this point. 

 

There is a lot of opposition to AI regulation, but there is a difference between open-ended regulation and the passage of specific laws that exact specific penalties for specific crimes.  Big tech firms are hard to punish compared to individuals, because their deep pockets make them treat fines as just another cost of doing business, and you can't send an entire corporation to jail.

 

The Gilbert and Sullivan operetta The Mikado has a famous ditty "My Object All Sublime," in which the emperor of Japan muses about how he's figured out punishments that fit the crime. 

Some ingenuity is needed here to come up with penalties for truly harmful hacking by escaped AI agents that would make AI firms highly motivated to prevent such breaches, either in the testing phase or after commercial sales. 

 

I don't know whether the following would be technically feasible.  But one suitable penalty would be the destruction of all copies of the AI model involved in the breach.  Models represent tremendous investments of time and money, plus the hopes of future gain.  So the destruction of the model responsible might hurt an AI firm more than any strictly financial penalty or sanction. 

 

Whatever we come up with to prevent these incidents in the future, it better work well, because in these relatively minor but significant recent breaches, we have received fair warning to do something about this problem before it causes serious harm.

 

Sources:  I referred to a report on NPR at https://www.npr.org/2026/08/08/nx-s1-5924878/meta-ai-breaches-external-firm-during-security-testing-sandbox-error and the Wired report by Lily Hay Newman at https://www.wired.com/story/openai-anthropic-ai-hacking-sprees-illegal/.