Two high-profile computer scientists from Princeton University, and coauthors of the book AI Snake Oil, have weighed in on the question of whether AI will kill us all by the end of the decade.
The pair have become well known for their work in separating fact from fiction in the AI world, and have now responded to an admission by Anthropic that there is a greater than 10% chance that AI could kill all humans within the next ten years …
AI killing us all within 10 years
Things kicked off when AI researcher Jacob Coxon, who has worked for both OpenAI and Anthropic, quit his job, stating that both companies were aware that AI systems could kill us all by the end of the decade. His post got more than 156 million views.
The people building Al earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately. No other human activity poses this level of danger.
You might expect the two companies to respond by dismissing him as a crackpot, but far from it. One of Anthropic’s AI safety leads, Evan Hubinger, said that Coxon was right, albeit over a slightly longer timeframe.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
There hasn’t been much specificity about exactly how AI might kill us, but the focus has been on the capacity for destructive cyber attacks on utility infrastructure (eg. water and electricity plants) and other key systems on which modern civilization depends.
Can alignment mitigate the risk?
A major element of the controversy is whether the risks can be successfully eliminated by more intensive work on alignment. This is the term used to mean ensuring that the objectives of AI systems are aligned with those of humanity.
Some argue that this will eventually eliminate the risk by ensuring that AI systems want the same things we do. Others disagree, saying that AI systems may pretend to be in alignment in order to pass safety checks while secretly holding conflicting objectives of their own.
Kapoor and Narayanan say it’s not enough
Sayash Kapoor and Arvind Narayanan have made a name for themselves by addressing what they call the “hype, misinformation, and misunderstanding” around AI.
The pair haven’t directly addressed the claim about the >10% risk of annihilation, but have written a lengthy essay in which they argue that it is possible to address AI risks through what they described as a layered approach to AI safety. They contrast the pessimistic view of the AI safety community with the more dismissive view of parts of the cybersecurity community that this is just business as usual.
They argue that a more measured middle ground is required. In particular, they say that alignment isn’t sufficient on its own but needs to be supplemented by three additional layers:
- Control, eg. sandboxing, human oversight, real-time monitoring
- Downstream defence, eg. organizations using AI to defend against AI attacks
- Resilience, eg. contingency plans for recovering from successful attacks
The pair end their essay on an optimistic note.
If we act with the appropriate urgency, we can hold AI companies to a higher standard, reduce risks from loss of control, and even tilt attacker-defender balance in cybersecurity back towards defenders.
9to5Mac’s Take
I don’t know whether AI will kill us by the end of the decade, but certainly there is enough reason for concern to make AI safety a key focus. The idea that alignment work is part of the solution but not sufficient in itself seems entirely sensible to me.
What’s your take on all this? Please share your thoughts in the comments.
Photo by Franck V. on Unsplash
FTC: We use income earning auto affiliate links. More.



