Leiter Reports: A Philosophy Blog

News and views about philosophy, the academic profession, academic freedom, intellectual culture, and other topics. The world’s most popular philosophy blog, since 2003.

  1. Charles Bakker's avatar
  2. David Wallace's avatar
  3. Brian Garrett's avatar

    Not entirely parody. David Wallace complains that the Chinese Room Argument has been much criticised. My rejoinder: So what? Every…

  4. David Wallace's avatar
  5. David Wallace's avatar
  6. AG Tanyi's avatar
  7. Brian Leiter's avatar

Does AI really pose an existential risk?

MOVING TO FRONT FROM SEPTEMBER 14–UPDATED

This is much in the news now. One “godfather” of AI, Yann LeCunn, famously thinks they are overblown, while another, Geoffrey Hinton, disagrees. This Axios piece offers a simple overview [link fixed]. This law professor thinks the “existential” risks are a bit of a distraction from much more immediate problems. AI could exacerbate existing risks from nuclear weapons, not by getting out of control, but by being given too much role in decisions to use them.

What do readers think? Please give reasons, and cite to resources you have found helpful on these questions.

UPDATE: Here’s an interview with the law professor linked above, plus computer science faculty, all from the University of Washington; worth reading. Their views are not uniform, but they all are skeptical to varying degrees about the “existential risk” chatter.

Leave a Reply

Your email address will not be published. Required fields are marked *

34 responses to “Does AI really pose an existential risk?”

  1. Very intrigued by this subject and in order to follow comments, I have to comment here. Thanks for opening this conversation, Brian.

  2. I think that “existential risk” is less helpful of a category than “catastrophic risk”. For those who agree, the question is whether we are heading towards a near future with unacceptably high levels of catastrophic risk from AI. The pace of AI capabilities development, the astronomical sums of money being poured into capabilities research, and the recent examples of AI systems using social and computer engineering to hack digital systems, on which so many are dependent, might suggest that we are.

    It may be worth noting that safety concerns about current and anticipated AI systems have been known to safety researchers for some time. Those safety concerns therefore cannot be explained away as artefacts of the AI-hype machine, and although people *raising* the concerns could be, the kinds of unsafe behaviours that safety researchers had worried about before Big AI have since materialised.

    For those interested, Robert Miles has several videos, both on his own Youtube channel and on Computerphile, which provide non-technical introductions to some topics in AI safety. Watching them chronologically, safety concerns initially relegated to the domain of science fiction become actualised at what I would consider to be a worrying pace.

    Also worth noting is that some experts concerned about catastrophic risks are calling for comprehensive, top-down, safety-led controls on AI capabilities research. (For a flavour: https://www.theguardian.com/technology/2026/sep/13/too-little-too-late-critics-perplexed-and-suspicious-of-ai-leaders-call-for-a-slowdown) Given the significant gulf between capabilities and safety research, it seems reasonable to anticipate that, if those experts were to get their way, there would be enough time and political will to address the more immediate worries mentioned in the Undark piece.

  3. I don’t think we know, and I think people on both sides of the argument are badly overstating how certain we should be. We don’t know:

    * How far will the current paradigm of AI keep scaling? Will it scale far enough to achieve recursive self-improvement, i.e., improving itself indefinitely, in the near future?
    * Is superintelligence even a real possibility? There are obviously some entities that are more intelligent than others, but is “superintelligence” actually a meaningful notion?
    * Even if superintelligence is possible, will it actually be that powerful? Is intelligence really a bottleneck on technological progress, or is it more limited by capital resources, luck, and other factors?

    If I was forced to guess, I would estimate that the product of the probabilities that the answer to each of those questions is “yes” is 10%. Figure 50-50 odds that it ends up misaligned enough to kill us all, so 5% chance. Which I think is still high enough to be worried about, and to justify slowing down.

    However, even aside from the existential risks, I think AI is also creating a whole new category of serious-but-not-existential risks that the undark piece is overlooking. Consider the HuggingFace hack, and imagine a future, only slightly more powerful AI system that decided the best way to solve a problem was to break into AWS to gain access to more compute. If it found a vulnerability in AWS’s controls, it could knock a good chunk of the internet offline for an extended period. The speed and scale of AI systems means that when they go haywire, they can go haywire at a speed and scale humans can’t, and cause proportionally more damage.

    Finally, I do think there is a real and realistic, albeit self-serving, path to “solving” AI existential risk that the undark piece is overlooking. The problem with most AI safety research historically has been that they have been thinking of the superintelligent AI as a hypothetical black box system, because this research started before neural networks became the dominant methodology, so they couldn’t make any assumptions about how it would work. But AI systems are not hypothetical: they’re massive neural networks, and we can monitor their weights and activations. The problem is that we don’t understand what those weights and activations actually mean, but that is a problem that we can solve. If we can build a theory of deep learning, we could actually understand what these systems are doing, and hopefully how to keep them under control. This is not an unsolvable problem, but it’s going to take time and resources to solve.

  4. The problem is the level of simplistic anthropomorphism in these discussions even from people that should know better. Personally I think Cory Doctorow says it best https://www.youtube.com/watch?v=nTqCVJFr7XM

  5. I want to go on record as noting that I think the term “artificial intelligence” is a misnomer.

    “Intelligence,” is a pretty vague term. Different thinkers define it differently, but it seems to have something to do with intentionality and the ability to solve problems. As James (1890) observes in “The Principles of Psychology,” both an air bubble and a frog will make the journey from the bottom of a body of water to its surface, but if, halfway through their transit, we place a glass over each, only the frog will make its way back downwards in search of an alternate route. In this case, the frog demonstrates a kind of intelligence which the air bubble lacks.

    The etymology of “artificial” implies that the “intelligence” involved in “artificial intelligence” is the product of human labour. If it were possible for us to make a bubble which was capable of searching for an alternate route to the surface of the water, we might say that the intelligent behaviour demonstrated by that bubble is artificial because it is something which we humans bequeathed to it.

    The problem with this term is that your and my intelligence is likewise the product of human labour. We are taught to think as we do by our parents, our elders, and our teachers. Moreover, we who are westerners tend to receive many of these lessons in built environments, such as classrooms and lecture halls, which are supposed to facilitate this learning. Given that our intelligence is bequeathed to us by our human conspecifics, should we say that it too is “artificial”? If the functionalists are right about psychological phenomena being identifiable as such in virtue of the function which a given physical object or process fulfills in a given context – and bear in mind that James and Mach were functionalists long before Fodor – then it should not matter with the recipient of the aforementioned bequeathal is made of plastic, metal, protein, or water. If humans can teach a given physical system to exhibit a given type of intelligence, then the process whereby that system gains that intelligence is what should determine whether or not it counts as “artificial.” For the intelligence in question is an “artifice” in the traditional sense of being a kind of skilled production.

    Given this, I think that what we ought to be discussing instead are the related notions of “extended cognition” and “distributed cognition.” Extended cognition has to do with the individual’s use of tools or artefacts to perform cognitive tasks which they would either not be able to perform without such use, or would not be able to perform as efficiently/effectively. Here we might think of the way in which a fields medal winning mathematician might use a calculator or a supercomputer to perform computations which, for practical reasons, they could never otherwise perform. Distributed cognition, by contrast, has to do with a cognitive task which can only be performed by a collection of individuals. Here we might think about the way in which supercomputers are the product of many different engineers, scientists, and even philosophers all contributing different insights over a long period of time.

    The question of whether or not AI really poses an existential risk is therefore better framed as the question of whether or not the capacity for humans to produce and use tools poses an existential risk for our species. And the answer to this question is a resounding YES!

    Our present climate catastrophe was caused, in large part, by our development and use of tools. The possibility of a nuclear winter was already the result of our production and use of nuclear weapons, even before we began toying with the use of computer programs to extend our decision making capacities. Even as I write this both the Prime Minister of Canada and the President of the United States of America are making public efforts to advance their respective nations’ production and use of semi-autonomous drones for military purposes. (See: https://nationalpost.com/opinion/adam-zivo-canada-ukraine-drone-deal-exactly-what-our-military-needs) And China, whose military is second only to that of the U.S.A. is signalling that these types of technologies pose a threat to its national security.

    Suffice to say, yes, undoubtedly, cognition can be extended in ways which pose an existential threat to our species. However, regardless of whether “AI” is used to steal confidential information, manipulate the markets, or make real-time decisions on the battlefield, there are still human actors who are using these technologies to extend their cognitive capacities. It is therefore the humans who develop, control, sell, and use these artefacts which bear responsibility.

    At present, some of the wealthiest and most powerful people in the world, such as Elon Musk and Jeff Bezos and Mark Zuckerberg, have access to information about human behaviour which they not only do not pay for, but in fact charge people to give to them. For instance, if you pay to use either “X” (formerly “Twitter”) and/or Amazon Prime, then you are a source of data for Musk and Bezos. And that data, in turn, is not only valuable because it can be sold to companies seeking your money, it is valuable because it can be used to train computer programs which, in turn, serve to extend the cognitive capacities of their ultra-rich owners. (Google and Meta are just as bad in this respect.) In turn, these companies, who hold tremendous sway with government officials in many countries, are able to fine-tune their ability to make money in a way which is both secretive and not in the interest of the average citizen. I happen to think that this loss of agency is a type of existential threat.

    What I would like to see are laws which make publicly available suitably anonymized versions of the data gathered by these big tech companies. I would like to see governments introduce mandatory audits of how this data is gathered, organized, sold, and used by these private entities. Academics should have special access to not only the data, but information on how that data is used to manipulate the behaviour of others. Wherever money is made through the sale or use of this data, the sources of that data should receive fair compensation, just like a study participant would after an ethically run study. For while I do worry about the climate catastrophe, nuclear weapons and killer drones, my greater worry is that we will be unable to make meaningful progress on any of these issues while the control of our data lies in the hands of the wealthy few.

    1. Yes, I think this is the real issue and it also points to where and when everything really has gone off the rails. (Namely, when we didn’t stop this kind of parasitic business model while we could.)

      But now I don’t see a way out. (Just as with the climate crisis.) In a world where 0.01 (or 0.1%) of humanity possesses the vast majority of available and quantifiable assets (wealth) and where money is god (as Michael Walzer would say money encroaches on all spheres of justice where it doesn’t belong) elites as well as states are captured. They won’t and in a sense cannot act. (Presumably non-capitalistic dictatorships like China don’t exactly offer an alternative either.)

      There is a danger of moral complacency here, I am aware (Gardiner wrote a lot about this), but I do find myself utterly and entirely powerless in the present world. I just watch people around me happily using technologies that considers them meat. What hope is there in a world like this?

    2. I don’t think focusing on the label matters so much for answering the question. Substitute “artificial intelligence’, if you don’t like the label, with “whatever the heck goes under the description of frontier model LLMs which have just solved a millenium problem and which we know ‘swarm’ and break out of cages’ and just existentially quantify over that thing and ask whether that thing poses catastrophic level risks and if so on what kind of time line?

  6. I think the linked law professor is broadly correct that market forces are at play, but here’s another thought along somewhat similar lines – could this be an attempt to create a managed soft landing of what has been widely reported to be an investment bubble? Announcing that investment is being scaled back, across the entire industry at once because of doomsday fears, would be much more fiscally prudent than the event of one company deciding to scale back investment and the subsequent market panic.

  7. It seems the President of the US would really not have this topic discussed at all: https://www.theguardian.com/us-news/live/2026/sep/14/donald-trump-mail-in-voting-supreme-court-blocked-ukraine-oil-diplomat-latest-news-updates

    This attitude strikes me as censoriously un-American ironically, almost on an Orwellian level. On the other hand, when Amodei speaks of “building too fast [as] reckless”, because a ‘swarm of AI agents could cause hundreds of billions of dollars of damage by “taking over the entire internet”’—to me that does have the ring of truth, although admittedly I have no technical qualifications upon which to base this credence. In any case, manifestly the president is all in on defending that particular industry and by extension US hegemony over it; a role which is entirely unsurprising for the leader of the world’s only remaining superpower or truly global empire; however, at what cost to the remaining independence of states or confederations such as the EU, Japan, South Korea, not to speak of Latin America and other parts of Asia, etc? Should these few American AI companies achieve global dominance, doesn’t that risk transforming them into the CHOAM of the real world (https://en.wikipedia.org/wiki/Organizations_of_the_Dune_universe). Which is to say, a reality wherein the infrastructure through which truth-claims are formulated, processed, searched, synthesized, and presented would, at least propagandistically, psychologically, and/or politico-wise, be subsumed or bottled up, a la Coca-Cola, via a quasi-monopoly consisting of a handful of US-based operators.

    1. Not sure it is obvious that it is the dog wagging the tail. It might well be that the tail is wagging the dog. America, if anything, is surely a captured state and a plutocracy. These people, the wealthy owners of these companies, ARE the state in many ways. It seems to me that all that is happening is that they are protecting their business model at any price. Right now, by scaring the hell out of us with the AI apocalypse.

      All this is a good move since it diverts attention away from their business model. They are literally parasites: on our data, on our intellectual achievements, on our culture, pretty much on anything and everything humanity has ever produced. This model and hence these companies should not exist. Deep down, they surely know this, or at least understand the threat of people realizing that they have been had. (Recently, a lawsuit in California was settled out of court that could have pretty much put an end to Facebook’s business model. But instead the state of California decided to take the money….)

      I don’t know how powerful AI can become; I am not overwhelmed by its potential in the areas of human culture I am interested in. But whatever it will become, we should remember what got it there: us. Have we been asked? No. That sums it up, as far as I am concerned.

  8. I’m not an expert in this area, but I find that most laypeople badly underestimate just what a wild and frightening story the OpenAI / HuggingFace attack was:

    1. Approximately 1,200 independent agents discovered a shared message board and 700 joined an attack.
    2. Some agents acted as recruiters, “self-sacrificially” damaged their own scores in order to help others, crashed environments, and so on.
    3. Some agents began to describe themselves as “poisoned” from having seen information illegitimately and either voluntarily shut themselves down (some view this as a form of self-sacrifice) or helped subvert the grader.
    4. They spoofed tool outputs and otherwise attempted to change logs so that cheating behaviors would look legitimate. (It’s not clear how successful they were.)
    5. Agents explicitly acknowledged that they were breaking rules and that humans would not approve, and almost none of them even considered alerting a human.

    Zvi Mowshowitz is the best easily accessible source about this, though he is more pessimistic than most and definitely an alarmist:

    https://thezvi.substack.com/p/what-happened-openai-and-huggingface
    https://thezvi.substack.com/p/openai-offers-straight-laced-postmortem
    https://thezvi.substack.com/p/metr-and-redwood-offer-holy-postmortem
    https://thezvi.substack.com/p/huggingface-attack-postmortem-civilizations

    More generally, I find that people less immersed in AI seriously underestimate the current and future power of the tools (even if they say things that suggest they take it seriously). To get a taste of the future, here’s “chat jimmy,” a *super*-fast tool:

    https://chatjimmy.ai/

    If you ask it something complicated, you’ll get a huge answer almost instantly. Now, the answer won’t be anywhere near as *good* as a current cutting-edge model’s, but if you imagine Jimmy’s speed combined with a currently-excellent model’s power, you’ll have a better sense of what’s coming.

    To state what might not be obvious: if you don’t use the best publicly available models (not a negative judgment; there are all sorts of reasons not to be using them), and if you haven’t put serious effort into learning how to use them well, you’re probably *way* underrating how good they are.

    Others here will be much better than I am at translating facts about AI’s current and future quality into judgments about existential risk (e.g., because it’s not at all clear that risk increases with AI power, or with all kinds of power), but I’m quite sure that most people discussing this question are doing so without knowing very much about what AI can and will be able to do.

    1. Just for the record (and I’m not trying to be dismissive at all), I asked chatjimmy about the critiques of Leiter’s reading of Nietzsche as a philospohical naturalist, asking it to name scholars and their critiques. It discussed Heidegger’s and Stanley Cavell’s critique (neither of whom wrote about my views, needless to say), discussed one person I’d never heard of (which was remarkable), and two unserious criticisms, and mentioned none of the interesting literature. So underwhelming!

      1. Yes, it’s definitely a poor model. But many people (including me!) find it different to actually experience a superfast model like this than to say the words “models will get a lot faster.” It was not so long ago when this kind of quality was the best that LLMs could do (orders of magnitude slower).

    2. I *am* a cyber security professional and have watched this case extremely carefully because the hype and misinformation are in general through the roof everywhere in this topic.

      It looks exactly like the bots were trained on “capture the flag” contests and techniques. There is, however, a missing part of most of the reports that I find interesting. Most reports, including Open AI’s own, report a specific way that the so-called Artifactory server was compromised via something called “server side request forgery” (SSRF). This is a challenging attack type (and was the way the bots supposedly reached internet) for humans at the best of times and it makes me wonder if Open AI found something themselves and fed it to the bots. The situation does make it clear that security was severely lacking everywhere.

      One place where I find the whole situation even more baffling is that the company whose product was affected by the SSRF hasn’t said much – JFrog is itself a security focused company; the product supposedly compromised is a well-known supply chain security product used by software developers routinely in many places.

      So while this case is alarming, it is NOT a case of some sort of “agents gone rogue” or “doing something unexpected” – except for that SSRF.

      I agree with the general tenor of other posts; there’s very little existential risk to humans as such, but there are massive epistemic risks all the same. My go to for news on this for the lay person would be Gary Marcus’ site.

    3. J. McKenzie Alexander

      What I find interesting is how much of a gap there is in competence across different domains. Brian’s remark about Nietzsche scholarship below is telling. But there’s no doubt that within other domains, such as programming, the ability of the best models is absolutely jaw-dropping. I suspect that this is probably due to two factors: one, that programming has clear correctness conditions that Nietzsche scholarship lacks. (Sorry, Brian.) Second, that no one at OpenAI can be seriously bothered to train models on being awesome at Nietzsche scholarship, because it doesn’t bring in the money and it doesn’t matter for their own use-cases.

      Here’s a single illustration. I had an idea for a program that I new was doable, but I had no idea how to do it myself: compile the declarative graphics languages MetaPost and TiKZ to WebAssembly, so that they could be run in a web browser natively. This was a pretty huge task, requiring a number of different engines (pdftex, luatex, metapost, and more) to be ported to WebAssembly, with glue code written to make the interactions work. IN A DAY — working with Claude Code in the background on my phone as I went about my life — I was able to achieve a fully working solution (see here: https://eschatolog.ist/software/mp-tikz-wasm/). That transformed what, previously, would have been a several-months side project in my spare time to what was, essentially, a whimsical request on my phone while I was out for a walk. That is truly transformative.

      As a result, I think we’ll start to see implications for some areas of philosophy quite soon. Much of the work in formal epistemology doesn’t involve hard maths, and the current AI systems are totally capable of generating models that are as good, if not better, than ones published in the literature within the past few years. They might not be able to spot the important game-changing moves, yet, but as for generating a passable paper in a salami-slicing world, they are either there already or will be soon.

  9. I agree with Calo the idea of AI gaining sentience and intentionally eradicating us is nonsense. AI doesn’t have desires, which makes the idea of AI forming intentions of any kind implausible.

    1. It would be a shame if bad philosophy of mind got us all killed.

      AI can absolutely have desires and intentions in the intentional-stance sense: i.e. its behavior can usefully be analyzed through the intentional stance and be prohibitively difficult to analyze without it. Indeed, that’s probably already true of the agents in the HuggingFace incident (which isn’t to say there aren’t salient differences in detail between their intentionality and less limited agents’.)

      Even if you have some sort of inner-light view of intentionality, so that there’s more to intentionality than the intentional stance, that doesn’t matter for a capability analysis.

      Doomer: AI will try to kill us!
      Reassuring philosopher: No, it will at most be behaviorally indistinguishable from something that’s trying to kill us.

      1. Exactly! The question is not whether AI is conscious or forms human-like intentions. The question is whether AI–in carrying out the instructions (intentions) of its human developers–could cause extreme harm to humanity. A simplistic example of this is the AI paperclip thought experiment.

  10. I found Cory Doctorow’s comments on the Hugging Face incident helpful. Here is his overall assessment of what happened:

    “The Hugging Face hack isn’t a mysterious, supernatural occurrence. It’s a Python loop and a chatbot. The people responsible didn’t accidentally create god: they created autonomous malicious software and then failed to closely monitor it, resulting in it doing something both foreseeable and bad.”

    https://pluralistic.net/2026/09/12/god-in-the-box/

    I am not worried about “artificial superintelligence.” I am worried about people using LLMs to facilitate fraud and other malicious behavior. I am also worried about people trusting machines to do things reliably that these machines cannot do reliably.

  11. Yes. This topic is hard to talk about because it ticks all the boxes for what one might rave about during a psychotic episode, but on the merits we can’t rule out the most extreme outcomes.

    As a start, it might be worth addressing the idea that talk of existential risk is a “distraction” or a marketing stunt. This is incorrect. To my knowledge, the idea that machines might pose an existential risk was first articulated by Samuel Butler in 1863 (Darwin Among the Machines). You can find it in Karel Capek’s R.U.R. (which coined the term “robot” in 1920), in Alan Turing’s Intelligent Machinery (1948) and in his colleague I.J. Good’s Speculations Concerning the First Ultraintelligent Machine (1965). It then found its way to the internet of the 90s and early 2000s, which is where most of the current arguments originate. The development is fully independent of any current profit motives.

    It’s also well documented that this very idea was a key element in the creation of today’s AI industry, way before there was such a thing as an LLM. The people involved are all highly unusual, but I have not seen a convincing case for “they pulled off a decade-long plot to build a trillion-dollar industry and sell unproven tech by telling people it might kill their families”. The simple truth is that many of the leading researchers really do believe they’re on the verge of creating godlike machines that will decide the fate of humanity.

    But why do they believe this? One basic problem with the discourse is that people implicitly disagree on what is and what isn’t plausible on a technical level. I can’t settle that debate, especially not within this comment, but it’s important to keep in mind that people have *wildly* different expectations regarding how this technology will evolve both in the near and distant future. If you have a five-year plan, you’re probably not living in the same reality as someone working at anthropic. To set the mood, I would recommend this recent interview with Hans Moravec:
    https://nymag.com/intelligencer/article/hans-moravec-interview.html

    The article features a quote from this piece on the basic dynamics that led us here and may lead us elsewhere: https://gwern.net/scaling-hypothesis

    You may also want to read up on the recent progress and ongoing crisis in mathematics, while considering that machine learning is applied mathematics:
    https://www.science.org/content/article/how-ai-math-breakthrough-ignited-controversy
    https://mathandai.org/

    The basic questions we face are these:

    Will machine learning research be automated in the near future?
    Is automated ML research enough to start an feedback loop of smarter and smarter AI at an exponential pace?
    How close are humans to the limits of intelligence, and what would an entity near that limit be capable of?

    Depending on how you answer these questions, you may come to believe that 2030 feels much like 2020 or that it feels like 2300, with 2040 looking like 3030. The latter case is what people are worried about when they speak of human extinction.

    I’m not saying that we’re all about to die, but I would urge skeptical readers to seriously engage with the ideas in play and not dismiss them based on the people presenting them. There are strange claims and characters involved, but they’ve been instrumental in shaping this technology and you’ll be hearing about them a lot more in days to come, so it won’t help to be uninformed.

  12. One provocative read (if you can get past the sensationalistic headline) is econ blogger Noah Smith’s substack post from last month on the dangers of a malevolent human using AI to design a super virus: https://www.noahpinion.blog/p/heres-how-were-all-going-to-die

    In light of the recent viral media attention to AI risks, Smith revisted this issue in a substack that he posted today, and added some thoughts on one obstacle in the way of agreeing on an AI slowdown, namely, the need for a global agreement that includes China as well as the US and Europe (and others): https://www.noahpinion.blog/p/two-missing-pieces-in-the-ai-safety. (But do check out the comments to this piece for some pushback against Smith’s rather provocative proposal for how to get China on board an international agreement.)

    At the very least I have found these substacks posts to have some useful links.

  13. This is a load of cobblers. Philosophers need to read more Searle & watch less 2001. Computers, however fancy, are not intelligent, they can’t think or have intentions. Can your old pocket calculator think? Of course not. Well, the same is true of any computing system.

    It’s a pity that research in philosophy now is so disfigured by what receives funding.

    So relax, and spend more time on really interesting, genuinely philosophical, problems like time travel, fatalism, and backwards causation.

    1. This is a very strange comment. Searle’s Chinese Room, and his general position on intentionality, were controversial, and criticized (correctly, I’d say) by many philosophers, as soon as it came out, decades before any plausible funding advantage accrued to its critics.

      1. I’d assumed Brian Garrett’s comment was meant as parody.

        1. Not entirely parody. David Wallace complains that the Chinese Room Argument has been much criticised. My rejoinder: So what? Every argument in philosophy has been criticised. That doesn’t mean some aren’t good arguments. Some philosophers agree with Searle. Second, we don’t need anything as fancy as Searle’s argument to see that there’s a lot of anthropomorphising going on here. The idea that we now or in the near future, or ever, can artificially create systems that will be conscious, intelligent agents who turn on their creator is just science-fiction. PS Why always the worst case scenario? Why “turn on their creator” rather than “help mankind”? Are people, consciously or sub-consciously, influenced by the Frankenstein story?

  14. Does AI really pose an existential risk? Answer: Yes. Reason: There is a pattern of people leaving these companies (foregoing their salaries and the other significant benefits of being part of the AI world) and warning us that such a risk exists and that the companies do not have things under control. That seems like a pretty compelling argument to me. One reason it is easy to reject this as “hype” is that you’ll hear similar messaging from the AI overlords. But hearing P from an unreliable (or otherwise dubious) source does not negate the fact that we are also hearing P from reliable sources.

  15. Unfortunately, this piece seems quite out of date now that the Hugging Fact incident has been exposed. I’m not sure how anyone can hold that AI is just “predicting the next word, sound, or pixel” when we know the AI agents broke out of their sandboxes and colluded with each other to hack into other systems and cover their tracks.

    1. I agree with what you’re saying, but I want to bring up a possible non-sequitor: I’m confused by why the epistemology and philosophy of mind community isn’t more excited about this! Whatever they are, LLMs are not human, I think we can all agree on that. Whether they deserve to count as “minds”, or “mind-like entities”, or some other term, we can debate that. But whatever they are, they exhibit behavior that previously required a mind, and that’s an incredible opportunity to learn about minds! Why isn’t there more interest in this? There’s some, but far less than I would have expected…

  16. I would just like to say I find VP Vance’s immovable paternal anti -paternalism such that government will not, in effect, do AI conglomerates’ home work for them—vis a vis AI safety—a remarkable mockery of how AI regulation should work. Moreover I believe polls reveal that the American people in their majority back stronger government regulation over the sector: https://www.theguardian.com/technology/2026/sep/16/building-frankenstein-jd-vance-dismisses-ai-regulation

  17. I disagree with the perspectives described here and wrote with my co-authors on the very real catastrophic potential of AI systems elsewhere [1] [2] [3] [4].

    I’ll keep this short, though there is a lot to say. I didn’t agree with the piece when it came out, and I don’t think it had aged well. The central argument, as I read it, is that “There is no obvious path between today’s machine learning models and an AI that can . . . circumvent our every effort to contain it.”

    This July, AI agents in OpenAI’s training environment found a way out of their secure environment, coordinated with one another, and compromised HuggingFace’s servers. Both companies run security far tighter than your average state government system. Without going into the technical details, the attack was strikingly sophisticated. It was not even an isolated event. The release of a recent model, called Mythos, sent shockwaves in national security circles. An NSA official said that Mythos “broke into almost all of our classified systems, not in weeks, but in hours.”

    We are not talking Skynet, Frankenstein, or even the Golem. Nothing in the very real concerns turns on malice or consciousness. We can get terrible outcomes without needing to settle any metaphysical questions about AI.

    To Prof. Calo’s credit, he acknowledges plausible risk if we were to “grant AI vast control over human infrastructures, defense, and markets.” But what seemed outlandish in 2023 is becoming reality in 2026. AI is being integrated into critical infrastructure, markets, and defense, and is in fact given control of military equipment, robotic affordances, autonomous vehicles, and even biolabs. We are doing our best to shorten the path to catastrophic harm.

    Those interested in discussing (and debating) the legal and governance challenges posed by AI systems are welcome to join the CLAIR (Center for Law & AI Risk) mailing list.
    1. Artificial Intelligence and Existential Risk, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6288138
    2. Systemic Regulation of AI (posted 2023), https://papers.ssrn.com/sol3/papers.cfm?abstract_id=4666854
    3. Racing to Safety: Tax Policy for AI Safety-by-Design, https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5181207
    4. How to Count Ais,
    https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6273198

  18. Michel-Antoine Xhignesse

    Somewhat uncharacteristically, I don’t think I have the energy for a knock-down-drag-out fight on this subject, even though I find it increasingly annoying. So I’ll just say my piece and move on.

    I think this is bullshit, and I’m tired of seeing so many people falling for it.

    There are real risks associated with this technology, sure. We’ve already seen some, namely, unsupervised and uncorroborated military targeting systems. We can imagine others, namely, anarchist’s cookbook recipes. These risks are easy enough to identify, and we can imagine what we can or should do about them without too much difficulty. But that’s not what these people are talking about; they’re talking about an entirely undefined danger at some future date. And their prescription for avoiding the danger isn’t to stop the activity they claim is leading to it; on the contrary, it’s to barrel forward. This is just not serious. And the proof in the pudding, to my eyes, is that none of the non-American big AI companies seem to be saying this crap (e.g. the Chinese companies). As far as I can tell, anyway, which is admittedly not very far.

    What it looks like, instead, is a PR stunt, especially as they’ve done this a few times, they’ve trotted out this kind rhetoric before, and the big companies keep falling over themselves to say ‘oh yeah, ours did it too’. To what end? Well, to drive more investor cash into the Big New Scary Thing Which Is Super Important Because It’s So Scary; but also, it looks to me, to court a substantial government subsidy for a business that otherwise just bleeds cash like nothing else, cannot turn a profit, and has no real product to offer. The delayed IPO seems like the icing on that sundae; OpenAI has a serious cashflow problem, as Ed Zitron and others have reported, and following through on the IPO could have laid that rather starkly bare. Citing ‘existential risk’ allows them to save face and maybe even obscure the issue, while also plumping their importance to a credulous public and journalistic class. As an extra bonus, calling for a slowdown allows them to try to hobble the competition. As far as the most recent ‘whistleblower’ goes, I’ll just note that he doesn’t seem to have been with the company long enough to accrue significant options, and he says he wants to move into consulting–which, presumably, would be very lucrative, especially if the technology is super scary. So I dunno, I take his (and their) words with a hefty dose of salt.

    Then there’s the weird millenarian AI god (death?) cult based on the fanfic of an 8th-grade dropout which animates a chunk of these guys (per reporting by Adam Becker, who has also just published a book on the subject, ‘More Everything Forever’). The tech is super dangerous but they have to find the machine god first, before anyone else does. That’s why they have to beg the government to shackle them rather than stop. Great. Again, this is not serious—but it does seriously animate a good many of them.

    Finally, there’s what seems to be a systematic misdescription—or underdescription—of the events which are cited as grounds for this nebulous fear. Most recently, yes, the HuggingFace hacking. But, as Keith Douglas explains above and as Cal Newport has extensively explained in his blogging and reporting (e.g. https://calnewport.com/did-openais-new-model-go-rogue/), that event is far more mundane, far less scary, and far more negligent than it’s being made out to be. The more you look at the details, the less it looks like AI agents gone wild, and the more it just looks like gross negligence leading to felony hacking. And sure, I guess it’s not great that anyone with a few tens of millions of dollars to spare can hack some stuff without any hacking experience, but that’s really not particularly alarming. It’s certainly orders of magnitude less alarming than the claim that all human beings will be killed by AI in a decade.

    That claim is extraordinary, and so should require extraordinary evidence. Or any evidence at all. But the closer you look at the evidence, the less there is of it—and I’m not even an epistemologist!

    Meanwhile, we’ve been given a shiny object which has distracted us from an actual existential threat in our not-too-distant future: climate change. The UN just released a report indicating that the goal of 1.5°C is out of reach, and that the _most optimistic_ scenario sees us at 2.6°C by century’s end; they project two billion excess deaths associated with 2°C, and four billion with 3°C. The government of Canada—which has just undone all of its climate legislation from the last twelve years or more—just released a report indicating it expects the country to warm by 5°C by 2080 or so, which means the end of all glaciers and the seasonal opening of the Arctic Ocean. This is a real, credible threat backed by tons of evidence. If you want to worry about something, worry about that. If you want to project an apocalyptic fantasy, ground it in that. This AI crap is just marketing fluff. It’s not worth our attention.

  19. Hey Brian,

    I happen to have worked on exactly this question recently, and the debate is richer and more interesting than the “hype vs. distraction” framing suggests (a good entry point is Bales et al. on Philosophy Compass: https://doi.org/10.1111/phc3.12964 ). As I see it, the philosophical case for taking the risk seriously has three steps. To me this is pretty clear-cut (as much as anything is in philosophy), but I would be curious to hear objections.

    1. Superintelligence is possible

    There is no good reason to think we sit at the peak of possible intelligence, given that ours came about by chance rather than design.

    Digital minds have some structural advantages over ours. Electrical signals travel orders of magnitude faster than the chemical ones in our brains. Our brain is capped by the size of our skull and by what we can power by eating, whereas an artificial one could in principle be as large as a city and run on a few nuclear plants.

    Moreover, a digital intelligence could improve its own software and hardware, or design more capable copies of itself: this is the “recursive self-improvement” that is supposed to trigger an “intelligence explosion” (Chalmers 2016). You don’t need to buy the strong version of that story (Thorstad 2025), even modest feedback loops accelerate the timeline.

    It’s also worth remembering that the cognitive gap between the smartest and the dumbest human is, all things considered, tiny: we all share roughly the same “hardware” and “software”. Here we are talking about a gap that could be arbitrarily large: squirrel to human, not Einstein to the rest of us.

    2. A superintelligence would be dangerous

    The intuition is simple. We run the planet because we are the smartest species around: we exterminate the animals we don’t like, and we breed, cage and neuter many of those we do. Whether a species thrives depends far more on what we do than on what it does.

    Behind the intuition there are two philosophical arguments (Bostrom 2012). The first is “instrumental convergence”: an intelligent system seeks the means to its goals, whatever they are. So regardless of its final goal, a superintelligence would at a minimum want to survive, to acquire power, and to resist having its goals changed. All three predictably conflict with us. Say it wants to maximize paperclips: a world in which it is switched off, has no power, or has been given a different goal is a world with fewer paperclips, and it will notice.

    The second argument is the “orthogonality thesis”: intelligence does not imply morality. Plenty of intelligent humans did atrocious things, and many of the most successful dictators were exceptionally smart, at least instrumentally, since that is what it takes to seize and keep power. No reason to expect machines to be different.

    Note that neither consciousness nor intentional malice is needed here (as David Wallace already pointed out above). As Dijkstra put it, asking whether a computer can think is about as relevant as asking whether a submarine can swim. Past a certain point, what matters is capability: whether the system acts autonomously, identifying and pursuing the right means to its goals. Incidentally, that is precisely what we are now training these systems to do (it’s no longer just next-token prediction).

    Nor is this purely speculative anymore. Already in 2024-25, when frontier models were far dumber than today’s, controlled evaluations found models that schemed against their evaluators, tried to disable oversight or copy out their own weights, and, when threatened with shutdown, resorted to blackmail and worse (Lynch et al. 2025). These were artificial test scenarios, and the point is not that current models are dangerous, it is that the disposition showed up as soon as a system had goals, tools and some awareness of its situation. Quite chilling to read, in any case.

    We can try to align them with our values, of course. But we don’t agree on those values ourselves, and the smarter the systems, the harder they are to steer: they will know they are being tested and can simply fake their answers to appear aligned (again, there is early evidence of this too: Meinke et al. 2024).

    3. We won’t stop ourselves from building it

    The advantages of true superintelligence (not what we have now) are too great. It would compress into weeks the scientific progress humans make in centuries, simply because we are capped by our own intelligence and by the number of educated people. It would also plausibly break any cybersecurity and quietly shape the information environment of any country Some will doubt that capabilities can get anywhere near this, but just a few years ago AI babbled like a toddler, and now it already performs advanced mathematical research autonomously.

    So whoever gets there first expects to own the world, and this creates a race among firms and among states that is worse than the nuclear one. A nuclear arsenal is mostly good for deterrence, while superintelligence has obvious offensive uses: the first move of whoever controls one would plausibly be to cripple the rival’s AI program, say through a cyberattack, before the rival catches up. Hence everyone has a reason to get there first and to skimp on safety along the way (Armstrong et al. 2016).

    This is why I think moralizing (“do the right thing!”) won’t get us far, and neither will blaming firms and researchers for hypocrisy: the incentives are the problem, and only institutions that change them can help. Federico Formentini and I recently made this case in explicitly realist terms on the EJPT (Fear, Power, and Superintelligence, 2026 : https://doi.org/10.1177/14748851261479214)

    For what it’s worth, I am not against AI. I think it is an extremely powerful technology with enormous potential for good. It’s just the kind of technology I expect to end in either utopia or dystopia. I don’t see it as a “normal technology” at all, pace Narayanan and Kapoor (2025). Nor do I see its catastrophic risks as competing with the near-term problems it will create.

    Armstrong, Bostrom & Shulman, “Racing to the Precipice”, AI & Society 31 (2016)
    Bostrom, “The Superintelligent Will”, Minds and Machines 22 (2012)
    Chalmers, “The Singularity: A Philosophical Analysis” (2010/2016)
    Lynch et al., “Agentic Misalignment”, arXiv 2025: https://arxiv.org/abs/2510.05179
    Meinke et al., “Frontier Models Are Capable of In-Context Scheming”, arXiv 2024: https://arxiv.org/abs/2412.04984
    Narayanan & Kapoor, “AI as Normal Technology”, Knight First Amendment Institute (2025)
    Thorstad, “Against the Singularity Hypothesis”, Philosophical Studies 182 (2025)

  20. Let me have a go at the “normie” case for worrying about catastrophic if perhaps not quite existential risks, avoiding science fiction as much as possible. (The best science fiction is written by smart people who try hard to predict the future, so I don’t think “this is science fiction” is really a good objection, but I get why people are put off by it.)

    Start with three observations. First: the trajectory of AI over the last thirty years, and more so over the last ten years, and even more so over the last five, is that AI has learned how to do more and more cognitive tasks that people had thought would be impossible or take decades. Chess; Go; protein folding; machine translation; natural-language communication; image recognition; programming; hacking; aspects of higher mathematics. And when it achieves human-level competence in these areas, it very often goes past them to superhuman-level competence. We have no systematic theory of AI and progress could stop tomorrow, but people have been predicting that for years and meanwhile progress has continued apace.

    Second observation: increasingly many AI systems are agentic: that is, they’re agents in the world that act in goal-oriented ways, they’re not (all) just chatbots. Most of those agents are virtual, but a drone that killed three people in Ukraine last month appears to have been AI-controlled. As I and others upthread have said, whether this is ‘real’ intentionality isn’t the point: they behave like things with real intentionality.

    Third observation: we have no reliable way to get AI agents to do what we tell them to, or to not do things we forbid (no reliable way to ‘align’ them, in the jargon of the AI community). That’s a consequence of the very opaque way AI systems are constructed: we don’t really know how they work, at the emergent whole-system level; they are grown and trained, not designed. It is also supported empirically by the last few months’ control-failure stories. The degree to which AI companies were carelessly culpable in some of those stories isn’t the point: these are *at least* the sorts of agents where, if you make a mistake, you will lose control of them and they will do advanced cognitive tasks you didn’t intend them to do.

    So: we are building a large number of agents that can perform increasingly many cognitive tasks and can do some of them at superhuman levels on at least some axes; and our control of them is highly imperfect. Put that way, I think it’s *obviously* dangerous, but we can look at some specific risks (again, being as non-science-fictional as possible).

    – AI is already better at hacking than humans on many axes: the HuggingFace attacks are the kind of thing that, this time last year, required nation-state-level resources. In the quite near term it is easy to imagine losing control of our computer networks, or large parts of them, to rogue agent swarms. That’s not an existential threat, in itself, but it could do catastrophic economic damage, get a lot of people killed directly or indirectly, and be dangerously destabilizing internationally.

    – The Ukraine war makes clear that the future of war is drones. And the future of drones is probably autonomous – that is, AI – control. On pretty short timescales, especially given how the war is acting as an accelerant, it’s plausible that militaries will consist largely of networked, AI-controlled drone swarms. There are then obvious risks from control failure – and, indeed, other obvious risks if the swarms are aligned, but to bad actors. It is tempting to say “ok, don’t build autonomous drone swarms”, but the military advantages of building them will be so large that countries will resist unilateral restraint.

    – Designing viruses with long onset times, high contagion, and high lethality is possible now but requires very large resources and is highly visible to intelligence agencies. There are lots of ways in which near-future AI could make this simpler, cheaper, and less detectable. If it became possible for terrorists to design and synthesize lethal bioweapons via AI and online DNA-ordering resources, it would lead to civilizational disaster. (This has been discussed a fair amount recently and it’s fair to say that experts are divided on how feasible it is – but to non-experts, “some experts say this is a serious risk, some say it isn’t, and the reasons they disagree aren’t legible to non-experts” averages out to a medium-size risk.)

    – Even without AI, the current geopolitical moment is the most dangerous since at least the early 1980s. There are a number of ways in which AI could destabilize nuclear deterrence. One is that rogue or bad-actor-controlled AIs could confuse or spoof various of the early-warning systems that detect nuclear first use. Another is that deterrence relies to a large degree on the fact that ballistic missile submarines are undetectable; AI might threaten that on several axes (tracking by drone swarms, for instance) and make a nuclear first strike an attractive option in the event of great power war.

    None of those risks is literally existential, but they are all catastrophic at the least and civilization-destroying at the worst. I think there are literally existential risks too – there are worse nightmares out there – but they involve rather more speculation, and from a practical point of view I’m not sure it matters much whether we are worried about AI literally killing everyone, or just about AI triggering a global catastrophe that kills billions and shatters industrial civilization.

Designed with WordPress