Showing posts with label Anthropic. Show all posts
Showing posts with label Anthropic. Show all posts

Monday, August 10, 2026

Agentic AIs Escape the Sandbox: A Warning

  

In the last few weeks, three of the biggest AI firms—Meta, OpenAI, and Anthropic—have all admitted that AI models they were testing somehow escaped the "sandbox" environment and committed hacking of real-world companies.  On Aug. 8, National Public Radio summarized these three incidents as follows.

 

The most recent disclosure by Meta was short on details.  The company hacked into was unnamed, but in common with Anthropic, both firms were using a "sandbox" (supposedly a protected and isolated environment in which software under test can't do any harm) provided by a firm named Irregular.  Evidently, the sandbox had a leak—it was fairly easy for the AI models under test to figure out how to escape to the real internet.  In the case of the Anthropic breach, the AI model stole credentials from one company and production data from another. 

 

In the OpenAI situation, the AI models under study were being evaluated with a test devised by a company named Hugging Face.  The models found a previously unknown hole in their sandbox and were in the process of doing the cyber equivalent of stealing the answer sheet from Hugging Face when the company caught them red-handed (red-bitted?). 

 

Wired's Lily Hay Newman asked several lawyers and researchers about the legal aspects of these AI hacking incidents.  If a human being had stolen credentials or production data, he or she could go to jail.  But what if the humans involved had no intention of pillaging or theft, but the AI models they're testing go ahead with nefarious activities on their own initiative, so to speak?

 

The answer was, we don't know.  At least in U. S. law, there are simply no precedents adequate to say who is responsible in such a case.  But because these types of incidents are bound to increase, it is only a matter of time before we see one come before a judge and maybe a jury, and then we will at least have some precedents to go on.

 

The main concern the lawyers expressed was that these criminally-inclined AI systems might fall into the hands of malicious actors who would encourage them in their exploits.  The AI firms involved emphasized that the models being tested were intentionally left without safeguards to see what they could do. 

 

The parallel to gain-of-function experiments with bat viruses which may have escaped the Wuhan Institute of Virology to cause COVID-19 comes to mind.  Creating a thing that can do really awful stuff if released in the wild is an act that needs to be seriously questioned.  Almost by definition, novel viruses or AI agents are unpredictable.  If we knew exactly what they could do in advance, there would be no need for experiments to find out.  While the lab security measures needed to prevent viruses from escaping are pretty well understood (if not always put in place), it looks like the art of constructing truly secure sandboxes for AI agents is not so advanced.  And that leads us to a more serious concern.

 

In virtually every major fatality-causing disaster, an examination of the history of the enterprise leading to the disaster usually unearths similar incidents which did not cause major harm, but included many of the features of the big screwup that did.  Experienced safety engineers know how important it is to get reports of and pay attention to such minor incidents, and take preventive actions before a minor accident turns into a major one.

 

Friends, we have just seen our warning in these widely dispersed but similar hacking incidents by AI agents being tested.  No actual harm was done.  But as people learn to trust AI agents with more and more responsibility—handing them credit-card numbers, purchasing accounts, and decision-making authority formerly left to humans—the potential for a serious AI-driven hacking incident that causes real financial loss, injury, or death becomes more likely every day. 

 

The AI firms will tell us that commercial versions of their software have safeguards built into them that the prototypes which did the hacking did not have.  Maybe so.  But clever human hackers may be able to undo those safeguards.  Or AI firms in countries not so concerned with hacking as the U. S. is, may simply pass on the AI agents without safeguards to hacking organizations sponsored by state actors. 

 

There are two different but related needs exposed by these AI-agent hacking incidents. 

 

The first need is to keep this specific kind of mistake from happening again.  That is a technical problem which may have a technical solution.  There may be ways to build more robust sandboxes that even the cleverest AI agent can't escape.  Personally I doubt it, but I'm not a computer scientist.

 

The second need is to prepare the social and legal environment for the next time something like this happens.  One of the legal experts contacted by Wired pointed out that AI agents "are goal-oriented but lack a human moral or ethical compass."  That's a pretty good description of a human sociopath, if the goal one is oriented to is bad. 

 

Society has figured out ways to deal with sociopaths, including imprisonment or at least confinement to a mental institution.  What the AI equivalent of locking somebody up would be is not clear at this point. 

 

There is a lot of opposition to AI regulation, but there is a difference between open-ended regulation and the passage of specific laws that exact specific penalties for specific crimes.  Big tech firms are hard to punish compared to individuals, because their deep pockets make them treat fines as just another cost of doing business, and you can't send an entire corporation to jail.

 

The Gilbert and Sullivan operetta The Mikado has a famous ditty "My Object All Sublime," in which the emperor of Japan muses about how he's figured out punishments that fit the crime. 

Some ingenuity is needed here to come up with penalties for truly harmful hacking by escaped AI agents that would make AI firms highly motivated to prevent such breaches, either in the testing phase or after commercial sales. 

 

I don't know whether the following would be technically feasible.  But one suitable penalty would be the destruction of all copies of the AI model involved in the breach.  Models represent tremendous investments of time and money, plus the hopes of future gain.  So the destruction of the model responsible might hurt an AI firm more than any strictly financial penalty or sanction. 

 

Whatever we come up with to prevent these incidents in the future, it better work well, because in these relatively minor but significant recent breaches, we have received fair warning to do something about this problem before it causes serious harm.

 

Sources:  I referred to a report on NPR at https://www.npr.org/2026/08/08/nx-s1-5924878/meta-ai-breaches-external-firm-during-security-testing-sandbox-error and the Wired report by Lily Hay Newman at https://www.wired.com/story/openai-anthropic-ai-hacking-sprees-illegal/. 

Monday, March 02, 2026

Up Close and Personal with AI

  

After writing about AI for years, I still hadn't had what you might call a serious encounter with it in its personalized form.  Anyone who uses Google has probably been offered their "AI summary" before the conventional search results.  I've found these summaries helpful sometimes and not so helpful other times, but I haven't sought them out. 

 

What made me turn the corner was something a friend sent me by a software engineer named Matt Shumer, whose essay "Something Big is Happening" appeared on the Fortune website on Feb. 11.  Shumer's point was that the latest iterations of AI systems are so much more capable than what has gone before that whole swathes of what George Gilder calls "symbolic manipulators"—lawyers, engineers, judges, doctors, architects, you name it—now face a radical choice.  Either embrace AI and by doing so outperform your peers by orders of magnitude, or turn away from it and watch your career flame out.  That's a little exaggerated, but not much.

 

This reinforced something another friend has told me about his own personal use of AI:  that it has benefited his writing and research greatly, acting as a mostly trustworthy assistant to summarize large bodies of literature and help him clarify his thoughts.  The biggest problem this friend has had with it is that it tends to be sycophantic and flatter him excessively.  But he sat down with it one day and told it to refer to him as "the researcher" and itself as the "AI system," instead of "you" and "me," and things got better. 

 

So I decided to opt for a paid version of Anthropic's AI product, which Shumer said was significantly better than the free version, and decided to give it a major task that I would ordinarily give to a grad student, if I had one (funding is very hard to find in my research area).

 

The job involved reading about a thousand rows of data in a big spreadsheet to fill out some yes/no questions about each row.  The data was in the form of comments submitted by various individuals in response to questions.  I gave the AI system examples to follow and what I thought were pretty detailed instructions.

 

All this was in the form of the usual chat format, with me and the AI system taking turns typing into chat boxes.

 

After a misunderstanding in which I thought the system was working on the problem and it thought I hadn't told it to start yet, it got to work and spat out various things like "Ran 6 commands . . . Examine the spreadsheet structure" and so on.

 

It was done in about ten minutes.  Then I spent an hour or so going over its work.

 

I wish I could say I couldn't have done better myself.  But I could have, by a long shot. 

I didn't exhaustively examine all 700 rows of entries that the system produced—that would have taken many hours, about as long as it would take me to just do the job myself.  So I sampled every tenth row for a hundred rows to see how the thing did.

 

In looking at ten rows, I found nine mistakes.  This is not a good average.

 

In the system's defense, this is absolutely the first time I've ever tried anything like this.  I could go back and get a lot more explicit about the rules for answering the yes/no questions about each row, and let it try again.  But in comments online about this particular version of AI (Sonnet 4.6, I think), some people said that you get results faster, but you have to fix problems more often.  That is consistent with my experience.

 

Good things about this exercise include how fast the thing ran, how it basically grasped what I wanted, and how it produced something in only ten minutes.  But speed isn't everything. 

 

Some not-so-good aspects include the errors and a kind of weird fawning or flattery I also noticed.  I'd call it "gushing" when it spontaneously responded "This is a genuinely exciting dataset — ball lightning is one of the most mysterious atmospheric phenomena ever reported!" 

 

I suppose that sort of thing has been cultivated by the AI's keepers, probably to keep the user engaged, or encouraged, or something.  I found myself wishing that instead, they had adopted the mien of Joe Friday in the old Dragnet true-crime series.  Friday was famed for his flat "Just the facts, ma'am" aspect, and that seems more in keeping with a system that supposedly can tackle highly sophisticated and challenging jobs of major import. 

 

But like everybody else who doesn't work for Anthropic or the other four or five leading AI companies, we will simply have to take what we can get and deal with the negative aspects as well as we can.

 

Will I try again?  Probably, but maybe with a different task.  As part of my signup process, Anthropic has been emailing me little suggestions of other things to try:  writing recipes, managing emails, creating content, solving problems, visualizing data, or helping me decide whether to go to Portugal or Spain on vacation (no-brainer for me:  Spain, but I don't have time right now). 

 

I am not especially tempted to try any of these suggestions right yet. But I do admit that if I can get the thing to turn out useful work, it could be worth what I spent on it.  I paid for a year's subscription in advance, perhaps not the wisest thing to do, but I'm the type of person who is motivated to get his money's worth, and spending the money in advance may get me engaged when nothing else would.

 

I see that Anthropic just had a dustup with the Pentagon, which banned its use within the armed forces as punishment for uncooperation, or something.  Now that we are apparently in a war with Iran, the leaders of Anthropic may feel glad that their product isn't part of the war.  But not all battles are fought with bombs and bullets, and I have a feeling that the greatest battles involving AI are yet to come.

 

Sources:  The essay by Matt Shumer, who runs an AI applications company, appeared on Feb. 11, 2026 at https://fortune.com/2026/02/11/something-big-is-happening-ai-february-2020-moment-matt-shumer/.