Showing posts with label OpenAI. Show all posts
Showing posts with label OpenAI. Show all posts

Monday, August 10, 2026

Agentic AIs Escape the Sandbox: A Warning

  

In the last few weeks, three of the biggest AI firms—Meta, OpenAI, and Anthropic—have all admitted that AI models they were testing somehow escaped the "sandbox" environment and committed hacking of real-world companies.  On Aug. 8, National Public Radio summarized these three incidents as follows.

 

The most recent disclosure by Meta was short on details.  The company hacked into was unnamed, but in common with Anthropic, both firms were using a "sandbox" (supposedly a protected and isolated environment in which software under test can't do any harm) provided by a firm named Irregular.  Evidently, the sandbox had a leak—it was fairly easy for the AI models under test to figure out how to escape to the real internet.  In the case of the Anthropic breach, the AI model stole credentials from one company and production data from another. 

 

In the OpenAI situation, the AI models under study were being evaluated with a test devised by a company named Hugging Face.  The models found a previously unknown hole in their sandbox and were in the process of doing the cyber equivalent of stealing the answer sheet from Hugging Face when the company caught them red-handed (red-bitted?). 

 

Wired's Lily Hay Newman asked several lawyers and researchers about the legal aspects of these AI hacking incidents.  If a human being had stolen credentials or production data, he or she could go to jail.  But what if the humans involved had no intention of pillaging or theft, but the AI models they're testing go ahead with nefarious activities on their own initiative, so to speak?

 

The answer was, we don't know.  At least in U. S. law, there are simply no precedents adequate to say who is responsible in such a case.  But because these types of incidents are bound to increase, it is only a matter of time before we see one come before a judge and maybe a jury, and then we will at least have some precedents to go on.

 

The main concern the lawyers expressed was that these criminally-inclined AI systems might fall into the hands of malicious actors who would encourage them in their exploits.  The AI firms involved emphasized that the models being tested were intentionally left without safeguards to see what they could do. 

 

The parallel to gain-of-function experiments with bat viruses which may have escaped the Wuhan Institute of Virology to cause COVID-19 comes to mind.  Creating a thing that can do really awful stuff if released in the wild is an act that needs to be seriously questioned.  Almost by definition, novel viruses or AI agents are unpredictable.  If we knew exactly what they could do in advance, there would be no need for experiments to find out.  While the lab security measures needed to prevent viruses from escaping are pretty well understood (if not always put in place), it looks like the art of constructing truly secure sandboxes for AI agents is not so advanced.  And that leads us to a more serious concern.

 

In virtually every major fatality-causing disaster, an examination of the history of the enterprise leading to the disaster usually unearths similar incidents which did not cause major harm, but included many of the features of the big screwup that did.  Experienced safety engineers know how important it is to get reports of and pay attention to such minor incidents, and take preventive actions before a minor accident turns into a major one.

 

Friends, we have just seen our warning in these widely dispersed but similar hacking incidents by AI agents being tested.  No actual harm was done.  But as people learn to trust AI agents with more and more responsibility—handing them credit-card numbers, purchasing accounts, and decision-making authority formerly left to humans—the potential for a serious AI-driven hacking incident that causes real financial loss, injury, or death becomes more likely every day. 

 

The AI firms will tell us that commercial versions of their software have safeguards built into them that the prototypes which did the hacking did not have.  Maybe so.  But clever human hackers may be able to undo those safeguards.  Or AI firms in countries not so concerned with hacking as the U. S. is, may simply pass on the AI agents without safeguards to hacking organizations sponsored by state actors. 

 

There are two different but related needs exposed by these AI-agent hacking incidents. 

 

The first need is to keep this specific kind of mistake from happening again.  That is a technical problem which may have a technical solution.  There may be ways to build more robust sandboxes that even the cleverest AI agent can't escape.  Personally I doubt it, but I'm not a computer scientist.

 

The second need is to prepare the social and legal environment for the next time something like this happens.  One of the legal experts contacted by Wired pointed out that AI agents "are goal-oriented but lack a human moral or ethical compass."  That's a pretty good description of a human sociopath, if the goal one is oriented to is bad. 

 

Society has figured out ways to deal with sociopaths, including imprisonment or at least confinement to a mental institution.  What the AI equivalent of locking somebody up would be is not clear at this point. 

 

There is a lot of opposition to AI regulation, but there is a difference between open-ended regulation and the passage of specific laws that exact specific penalties for specific crimes.  Big tech firms are hard to punish compared to individuals, because their deep pockets make them treat fines as just another cost of doing business, and you can't send an entire corporation to jail.

 

The Gilbert and Sullivan operetta The Mikado has a famous ditty "My Object All Sublime," in which the emperor of Japan muses about how he's figured out punishments that fit the crime. 

Some ingenuity is needed here to come up with penalties for truly harmful hacking by escaped AI agents that would make AI firms highly motivated to prevent such breaches, either in the testing phase or after commercial sales. 

 

I don't know whether the following would be technically feasible.  But one suitable penalty would be the destruction of all copies of the AI model involved in the breach.  Models represent tremendous investments of time and money, plus the hopes of future gain.  So the destruction of the model responsible might hurt an AI firm more than any strictly financial penalty or sanction. 

 

Whatever we come up with to prevent these incidents in the future, it better work well, because in these relatively minor but significant recent breaches, we have received fair warning to do something about this problem before it causes serious harm.

 

Sources:  I referred to a report on NPR at https://www.npr.org/2026/08/08/nx-s1-5924878/meta-ai-breaches-external-firm-during-security-testing-sandbox-error and the Wired report by Lily Hay Newman at https://www.wired.com/story/openai-anthropic-ai-hacking-sprees-illegal/. 

Monday, March 20, 2023

Trying Out ChatGPT

 

Since the artificial-intelligence laboratory OpenAI made its latest major project ChatGPT available to the public last fall, the chatbot's popularity, not to say notoriety, has soared.  Chatbots—software that responds to human-typed inputs with conversation-like output—are nothing new, but the combination of speed, apparent knowledge, and polish with which ChatGPT responds to a huge variety of "prompts"—basically, commands to write something about a subject—have attracted probably millions of users, a ton of publicity, and expressions of concern.

 

One of the most understandable concerns is that students will simply take any given writing assignment, put it into ChatGPT, and cut-and-paste the result into their homework.  Plagiarism is a chronic problem in education, and universities across the world have been holding special meetings to deal with the advent of ChatGPT and how to detect and prevent such cheating. 

 

I wish them luck, because when I tried the system this morning on a topic that's very familiar to me, it came up with verbiage of such high quality that I wouldn't hesitate to use it as the lead section of a research proposal, for instance.  That is, if I didn't mind the fact that I was using some computer's synthetic prose rather than my own. 

 

In case you want to judge for yourself how ChatGPT did, here's a sample.  The prompt I gave it was this:  "Describe ball lightning in two paragraphs or less (under 250 words) and quote experts in the field." 

 

The response begins, "Ball lightning is a rare and mysterious phenomenon in which a glowing sphere of light appears during thunderstorms and floats through the air for several seconds to several minutes before disappearing."  So far, so good.  It goes on for 136 words, which is under 250, and quotes only one expert, John Abrahamson.  The quote itself is a long one—57 words—and seems to be taken from an interview that I was not immediately able to identify by typing it into Google, a favorite trick I used to pull with student essays that I suspected of being copied wholesale from the Internet.  Either Google doesn't do that type of search very well anymore, or ChatGPT may have used some obscure transcription of a radio or TV interview, but not even part of the original two sentences shows up in my search.  So I simply have to take ChatGPT's word for it that it's accurately quoting Prof. Abrahamson, a New Zealand chemical engineer who published a well-publicized theory of ball lightning around 2000.

 

And that points out one of the big problems with some forms of AI:  they behave like black boxes, and figuring out how they work and where they get their information can be difficult or simply impossible.  I suppose I could go back and ask ChatGPT where it got the quote, but then I wouldn't have time to finish this column. 

 

So is access to powerful software such as ChatGPT a threat to the integrity of education and the livelihood of copywriters and grant writers everywhere, or on the other hand a great boon to the millions of people who can't put two coherent sentences together?  To some extent, I'd have to say "all of the above." 

 

Whatever else the ChatGPT developers have done, I have to congratulate them on the generally flawless grammar in all ChatGPT outputs I've seen so far.  They must have come up with some way of assessing the grammatical quality of sources and picking only the best ones, because believe me, there is a lot of bad English grammar out there, especially in the reams of technical publications that attract authors whose first language isn't English.  So that's the good news.

 

What is perhaps not so good news is that lots of us could become dependent on ChatGPT and its successors.  Now, is this a dependence that is harmless, like our dependence on pocket calculators instead of doing long division by hand?  Or is it a malignant dependence such as some people have for porn or alcohol or video games, distorting their lives and inhibiting human flourishing? 

 

My first impression is that the main hazard so far of using ChatGPT is that of letting the machine do one's writing and thinking too.  Now, technically, I let my pocket calculator do my thinking when I use it, but the kind of thinking it does is extremely mechanical—that's why mechanical calculators were successful—and it's no loss to my mental integrity to outsource the taking of square roots to a machine. 

 

But expressing a complicated original idea in clear prose is something that has thus far been reserved for humans.  If I take out the word "original," it appears that ChatGPT can do as good as or better than your average human being at expressing complicated ideas clearly.  And of course, original is a relative term, as nobody can come up with a fourth primary color, for example.  We quickly get into philosophical waters here, but I will leave it with the Christian observation that God is the only Person who can truly originate things from nothing.  All so-called human inventions and discoveries are the unearthing or understanding of things and ideas that have always been latent in the universe, waiting for us to find them. 

 

I don't know whether some puckish mathematician has yet typed into ChatGPT, "Prove Goldbach's conjecture true or false."  Goldbach's conjecture is the proposition that every even number greater than 2 is the sum of two primes.  It's one of those things that seems like it ought to be true, and nobody can find a counterexample, but nobody so far has been able to prove it one way or the other.  From everything I understand about ChatGPT, it would come up with a lot of verbiage, and maybe equations, but as it simply pulls from whatever is already out there on the Internet (and according to its developers, it's skimpy on anything after 2021), if a proof isn't out there it's not likely to come up with one.

 

So the mathematicians are safe.  For the rest of us, I'm not so sure.

 

Sources:  A good description of ChatGPT and instructions on how to use it were published on the website Digital Trends at https://www.digitaltrends.com/computing/how-to-use-openai-chatgpt-text-generation-chatbot/.  I also referred to a list of ten hardest unsolved math problems at https://www.popularmechanics.com/science/math/g29251596/impossible-math-problems/ (You don't think I really go around worrying about Goldbach's conjecture, do you?)

Sunday, August 14, 2022

AI Illustrator Still Looking for Work: The Shortcomings of DALL-E 2

 

The field of artificial intelligence has made great strides in the last couple of decades, and any time a new AI breakthrough is announced, critics voice concerns that yet another field of human endeavor has fallen victim to automation and will disappear from the earth when machines replace the people who do it now. 

 

Until recently, the occupation of art illustrator seemed reasonably safe from assault by AI innovations.  Only a human, it seemed, can start with a set of verbal ordinary-language instructions and come up with a finished work of art that fulfills those instructions.  And that was mostly true until AI folks began tackling that problem.

 

The AI research lab OpenAI has publicized its work in this area using a "transformer" type of program that has gone through two versions so far, DALL-E and now DALL-E 2.  I'm sure the late surrealist artist Salvador Dali would be pleased at the honor of having a robot artist named after him, but the connection between his works and the productions of DALL-E 2 are perhaps closer than the researchers would like.  In an article in IEEE Spectrum, journalist Eliza Strickland highlights the shortcomings of DALL-E 2's productions and speculates on whether such systems will ever be generally useful.

 

As is the case with most such human-task-imitating systems, the first step the researchers took was to compile a large number of examples (650 million in the case of DALL-E 2) to train the software about what illustrations look like.  The images and their accompanying descriptive texts were all from the Internet, which has its own biases, of course.  The researchers have learned some lessons from previous fiascos with AI software that allowed random users to request sketchy or offensive products, so they have been very careful about pre-screening the training images and (in some cases) manually censoring what DALL-E 2 comes up with.  They have also not released the program for general use yet, but have carefully selected users under controlled conditions.  If you just let your imagination roam with the scenario of letting some randy teenage boys loose with a program that will make a picture of whatever they describe to it, you can see the potential for abuse.

 

For certain purposes, DALL-E 2 does fine.  If you want a generic type of picture that would look good as a filler for a brochure about a meeting, and you just want to show some people in a corporate setting, DALL-E 2 can do that.  But so can tons of free clip-art websites.  If you want an image that conveys specific information, however—a diagram, say, or even text—DALL-E 2 tends to fall flat on its digital face.  The Spectrum journalist was privileged to make a few text requests to DALL-E 2 for specific images.  She asked for an image of "a technology journalist writing an article about a new AI system that can create remarkable and strange images."  She got back three photo-like pictures, but they were all of guys, and only one seemed to have anything to do with AI.  Then she asked for "an illustration of the solar system, drawn to scale."  All three results had a sun-like thing somewhere, and planet-like things, and white circular lines on a black background showing the orbits, but the number of planets and what they looked like was pretty random.  So technical illustrators (those who are left after software like Adobe Illustrator has made every man or woman his own illustrator) need not file for unemployment insurance right away.  DALL-E 2 has a way to go yet.

 

There is a fundamental question lurking in the background of AI exploits like this.  It can be phrased a number of ways, but it basically amounts to this:  will AI software ever show something that amounts to human-like general intelligence?  And believe it or not, your humble scribe, along with Gyula Klima, a philosopher then at Fordham University, recently published a paper addressing just that question, and we concluded that the answer was "No."

 

As you might guess when a philosopher gets involved, the details are somewhat complicated.  But we began with the notion that the intellect, which is a specific power of the human mind, relies on the use of concepts.  In the limited space I have, I can best illustrate concepts with examples.  The specific house I live in is a particular thing.  There is only one house exactly like mine.  I can remember it, I can form a mental image of it, and I can even imagine it with a different color of trim than it actually has.  And software programs can do what amounts to these sorts of mental operations as well.  In my mind, my house is a perception of a real, individual thing.

 

By contrast, take the concept of "house."  Not my house or your house, just "house."  Any mental image you have that is inspired by "house" is not identical to "house"—it's only an example of it.  The idea or concept denoted by the word "house" is not reducible to anything specific.  The same goes for ideas such as freedom or conservatism.  You can't draw a picture of conservatism, but you can draw a picture of a particular conservative. 

 

In our paper, we gave strong evidence in favor of the notion that because concepts cannot be reduced to representations of individual things, AI programs will never be able to use them.  In the article, we used examples from an art-generating AI program that in some ways resembles DALL-E 2, in that it was trained on thousands of artworks and then made to generate artwork-like images.  The results showed the same kind of fidelity to superficial details and total absence of underlying coherence that DALL-E 2's productions showed.  As one of the OpenAI researchers quoted in the Spectrum article noted, "DALL-E doesn't know what science is . . . . [S]o it tries to make up something that's visually similar without understanding the meaning."  That is, without having any concept of what it's doing.

 

The big question is whether further research in AI will produce programs that truly understand concepts, and use that understanding to guide their production of art, text, or what have you.  Klima and I think not, and you can look up our article to understand why.  But we may be wrong, and only time and more AI research will tell.

 

Sources:  Eliza Strickland's article "DALL-E 2's Failures Reveal the Limits of AI" appeared on pp. 5-7 of the August 2022 print issue of IEEE Spectrum.  "'Artificial intelligence and its natural limits" by Karl D. Stephan and Gyula Klima appeared in vol. 36, no. 1, pp. 9-18, of AI & Society in 2021.