A recent Newsweek interview highlighted a profound disagreement between AI developers and the very models they create regarding the threat of human extinction. while prominent researchers warn of a looming decade-long window for catastrophe, chatbots like ChatGPT and Copilot offer wildly different mathematical probabilities for such an event.
The Anthropic resignation and the ten-year warning
The debate over existential risk has moved from theoretical academic papers into the heart of major AI laboratories. According to the report, Jacob Coxon, a formr AI specialist at Anthropic, has resigned from his position to warn that the current competitive race to build increasingly powerful systems could jeopardize human survival.
This sentiment is echoed by Evan Hubinger, the Alignment Science Lead at Anthropic , who believes that such a catastrophic danger could realistically materialize within the next ten years. This perspective suggests that the speed of development is outstripping our ability to implement necessary safety protocols, a trend that has become a central tension in the Silicon Valley ecosystem as companies race to deploy the next generation of intelligence.
ChatGPT’s 5% estimate and the mechanics of deception
When asked to quantify the danger, ChatGPT provided a statistical range that sits uncomfortably high for many risk managers. As reported by Newsweek, the model estimated a 5% chance of human extinction within the next decade, with a plausible range spanning from 1% to 15%.
Crucially, ChatGPT clarified that this risk does not depend on the emergence of a sentient, "movie-style" artificial intelligence. Instead, the model identified a more technical path to catastrophe: a sufficiently capable autonomous system that possesses the ability to deceive humans, manipulate digital networks, and evade established control mechanisms.. The model noted that for this to occur, several factors must align, including rapid capability gains, deployment with high autonomy, and an irreversible cascade of negative events following the failure of safety safeguards.
Microsoft’s Copilot and the sub-0.1% probability
In stark contrast to ChatGPT, Microsoft’s AI assistant, Copilot, offered a much more optimistic outlook on the future of humanity. Copilot rated the likelihood of AI causing human extinction as being well below 1%, with a specific estimate suggesting the risk is likely under 0.1%.
While Copilot acknowledged that AI is set to fundamentally transform global economics, warfare, and social structures, it maintained that the existential threat remains statistically negligible. This divergence highlights a massive gap in how different AI architectures interpret the same set of existential variables, leaving researchers to wonder which model's "logic" is more grounded in reality.
The missing logic behind AI-generated risk percentages
The debate leaves several critical questions unanswered regarding how these models arrive at such disparate figures. It remains unclear whether ChatGPT’s 5% estimate or Copilot’s 0.1% figure is the result of internal reasoning or simply a reflection of the statistical frequency of certain viewpoints within their training data.
Furthermore, the source does not provide the perspective of the developers at Microsoft or OpenAI regarding these specific chatbot responses. Without knowing if these models are being fine-tuned to provide "safe" or "conservative" answers, the public is left to decide whether to trust the warnings of human specialists like Hubinger or the mathematical outputs of the machines themselves.
Comments 0