Monday, September 21, 2026

Should We Run For the Hills From AI?

  

In the last few weeks, at least two well-publicized resignations from AI firm Anthropic have drawn attention to the pace at which AI models are getting more powerful all the time.  Anthropic's safety head Mrinank Sharma has stated his intention to resign as of next Feb. 8 in an open letter.  And Jacob Coxon has stated in an X thread that "the people building AI earnestly believe it could kill us all by the end of the decade."  According to Wikipedia, Anthropic's leading safety researcher, Evan Hubinger, says that there is a greater than 10% chance that AI systems could "kill all humans" in the next ten years.

 

Reactions to these outcries have ranged from indifference to panic or cynicism, with some noting that Anthropic is planning an initial public offering of stock this coming November.  Recalling the adage that there's no such thing as bad publicity, meaning that just getting your name out there is more important than what they're saying about you, this last take may have a grain of truth in it.  But the more likely explanation for this spate of criticisms and resignations is that something is going on in AI firms that scares the socks off some of those working on it.

 

Predictions of superintelligent systems taking over the world are not new.  I have on my shelf a book by Oxford philosopher Nick Bostrom called Superintelligence that was published in 2014.  In it he examines in thorough detail the various ways that AI could break bad, and discusses possible countermeasures humanity can take.  After three hundred pages of analysis that examines most if not all of the disaster scenarios that people are discussing today, he concludes]the book with these words:  "In this situation, any feeling of gee-wiz [sic] exhilaration would be out of place.  Consternation and fear would be closer to the mark; but the most appropriate attitude may be a bitter determination to be as competent as we can, much as if we were preparing for a difficult exam that will either realize our dreams or obliterate them."

 

The main fear is that AI will get out of control in some harmful way.  That has already happened with the hacking of the cloud-computing company Hugging Face by OpenAI models that were supposed to be isolated in a "sandbox," but managed to escape and damage a good portion of Hugging Face's infrastructure.  This is the most well-publicized example of an AI system that did something bad that its supposed controllers did not intend.  OpenAI says that it's taken measures to prevent such incidents in the future, but as we have pointed out elsewhere, this incident looks a lot like a warning that worse things could happen if AI firms don't implement vastly more secure means to keep experimental AI models under control.

 

The leaders of several AI firms in the U. S. have considered a mutual arrangement to slow down the pace of improvements to allow security measures to keep up.  This has caused concerns that China will catch up to or surpass U. S. competence in AI.  Discussions of government regulation have been discouraged by President Trump, who seems to have never met an AI model he didn't like (never mind about other kinds of models).  So if the government isn't going to do anything significant, and the companies can't get together on how to moderate the pace of improvement, what should the average citizen do?

 

Harking back to an existential threat that I experienced when I was in my twenties, the threat of nuclear annihilation was ever-present as I was growing up.  Nobody I knew developed nuclear weapons, but we all had to live with the consequences of the two superpowers being able to vaporize each other in a matter of a few hours.  There's an existential threat for you.  Somehow (my own hypothesis is it was the grace of God), we have managed to get from then to now without ending it all in a forest of mushroom clouds. 

 

AI is a different kind of threat, admittedly.  The kinds of systems that are most vulnerable to it seem to partake of the nature of AI itself:  large-scale networked systems that are centrally controlled.  There are certain infrastructure vulnerabilities in the U. S. that genuinely concern me, notably the fragile nature of  power grids.  They rely on substation transformers that cost upwards of millions of dollars, are hugely inconvenient to repair or replace, and are very limited in supply.  There have been news reports recently about how a concentrated and coordinated attack on substations nationwide—either physical, which would be hard to do, or a cyberattack, which might be a lot easier—could cripple the grid enough to leave large swathes of the U. S. without electric power for months.  That could lead to a civilization-destabilizing crisis.

 

Fortunately, there are some things we can do about this.  One is ramping up production of spare transformers (preferably making them in the U. S.) so that we have enough to replace as many as we might lose in a widespread attack.  Another is to take advantage of smart-grid capabilities now coming online to make it possible to operate pieces of the grid without having it all interlocked together.  As more renewable sources come online that are distributed rather than coming in a few big lumps like traditional power plants, one can imagine making the power grid more like the internet, which was designed to be robust against damage to a good many of its nodes. 

 

I'm not a power-systems engineer and this may be a fantasy, but it seems like it's within the realm of possibility.  Of course, determined hackers can wreck even well-protected infrastructure.  But making the grid more local, which is what we're really talking about, makes it harder for an attacker to take down all of it.  If local pieces have to be dealt with one by one, there is safety in variety.

 

I hope we can keep AI under control to the extent that the benefits outweigh the liabilities.  But if the history of engineering ethics is any guide, though, we will learn how to do this by dealing with the aftermath of some failures first.

 

Sources:  I referred to posts by Mrinank Sharma at https://www.linkedin.com/posts/andrew-hilger-436b713_the-attached-resignation-letter-from-mrinank-activity-7427113299336454144-4CVm/and Jacob Coxon at https://x.com/hilbertspaess/status/2097476196791709843 as well as the Wikipedia articles on "Existential risk from artificial intelligence" and "OpenAI-Hugging Face incident."

 

Update on Reader Survey:  As I suspected, there appear to be relatively few of you who read this blog regularly.  But I am very grateful to all nine who responded to the survey, and hope there are a few more who didn't bother to respond but still read it regularly.  You are in exclusive company!

 

No comments:

Post a Comment