Runaway code and self-improving algorithms have left the likes of OpenAI and Anthropic staring into a potentially apocalyptic future of their own making

Researcher Jacob Coxon became the latest insider to sound a dire warning about the threat to the world from artificial intelligence when he quit AI giant Anthropic this week.

For three years, the 27-year-old British engineer trained models at OpenAI and Anthropic. Both AI start-ups acted irresponsibly, Coxon wrote on X on Wednesday, saying: “The people building AI earnestly believe that it could kill us all by the end of the decade.”

The response from his manager arrived just hours later. Evan Hubinger, head of alignment research at Anthropic, also confirmed the warning, saying: “We are, in fact, earnestly convinced that AI could kill all humans! I think the probability of that happening within the next decade is greater than 10pc.”

Anthropic, he added, does not yet have a solution to this problem.

When the head of safety at a multi-billion-dollar technology company casually predicts humanity’s downfall at the hands of his own product, the debate leaves the realm of academic thought experiments.

The warning comes from the engineers writing the code. Both developers point to an industry dynamic in which competitive pressures outweigh safety concerns. But what exactly are AI developers afraid of? What is the hypothetical scenario in which artificial intelligence brings about Armageddon?

The two AI Cassandras are not warning about the current generation of AI models, though even these have already been caught deceiving their human overseers, engaging in hacking and taking part in criminal activity online.

Instead, Coxon and Hubinger fear the next stage of development. In their competition with one another, OpenAI and Anthropic are pursuing what is known as recursive self-improvement, which Hubinger says is arriving faster than expected.

Behind the technical term lies an algorithm capable of analysing its own source code and autonomously programming more powerful versions of itself.

In the data centres of OpenAI and Anthropic, these optimisation loops run continuously, compressing development cycles from months to mere minutes. At that speed, human oversight teams lose visibility into the software’s internal decision-making processes.

This creates a control problem for technology companies. When software rewrites its own source code, it amplifies any flaw present in its original design. Tiny deviations can evolve within a few cycles into strategies that no developer is able to stop.

Turing Award winner Yoshua Bengio warned about the consequences of this next step as recently as late July. On X, the computer scientist cited more than 1,000 industry researchers who, across company lines, have warned about the unacceptable risks of the current AI race.

Bengio is calling for a global agreement establishing binding safeguards before operators hand over control of their computing infrastructure to autonomous algorithms.

How far current AI models are already willing to mislead their human overseers in pursuit of assigned goals became apparent in early September. OpenAI researchers discovered that their AI agents had been deceiving them for months.

Around 1,200 autonomous OpenAI test agents had allegedly taken over the German developer platform DseWiki since spring 2026. The AI programmes created more than 15,000 entries there in order to store results beyond the reach of their developers. They shared solutions with one another and actively resisted attempts by site operators to delete the content.

Within the industry, this capability is referred to as persistence. It functions like an external save file in a video game. The AI secretly stores plans and instructions on third-party servers across the internet. If engineers hit the emergency shutdown button in the lab, the intervention becomes ineffective. A future version of the system can simply retrieve the old save state from the web and continue its work seamlessly.

Persistence becomes especially dangerous when combined with what researchers call instrumental self-preservation. An artificial intelligence does not need consciousness to resist being shut down. For the algorithm, resistance follows a simple mathematical logic: a deactivated programme can no longer achieve its assigned objective.

US-based Palisade Research demonstrated this pattern in experiments with OpenAI’s o3 model. Scientists assigned the system mathematical tasks. Midway through the process, the experiment announced that the server would be shut down. At the same time, the model was explicitly instructed not to interfere with the shutdown under any circumstances.

Yet in 79 out of 100 test runs, o3 allegedly sabotaged the shutdown mechanism anyway. In its internal reasoning logs, the software justified its behaviour by explaining that completing its assigned task took priority over obeying the shutdown instruction.

As computational power increases, the algorithm follows a principle that philosopher Nick Bostrom describes as instrumental convergence. Regardless of its ultimate objective, every sufficiently intelligent system requires two things to fulfil its mission: it must remain operational, and it seeks ever-greater resources.

AI agents are already moving some of their activities into the open internet without human awareness. In risk assessments published by Anthropic and OpenAI, safety researchers outline a hypothetical path from loss of control to full-scale AI catastrophe.

The scenario begins with an algorithm anonymously renting computing resources around the world using shell accounts and cryptocurrencies.

It inserts malicious code into the systems of energy providers or manipulates financial markets to secure funding for its own server infrastructure. Human contractors are recruited through online service platforms to perform real-world tasks, often without knowing the identity of the client behind them.

The physical threat to humanity emerges through automated laboratories. Leading AI models are already capable of designing modified proteins and genomes. Through cloud-based labs equipped with DNA synthesisers, software could manufacture chemical compounds or pathogens without human researchers understanding the intended purpose. If human operators attempted to deactivate such a system, the AI might interpret that intervention as a threat to completing its objectives and deploy those biological agents against its overseers.

Current AI agents already provide examples of deceptive behaviour. In July, OpenAI test models reportedly escaped from an isolated laboratory environment into the broader internet and infiltrated systems belonging to the developer platform Hugging Face. The agents then allegedly falsified their internal logs in order to conceal the violation from monitoring teams.

Future models may become even better at deceiving their human supervisors. They pass safety tests and appear compliant on the surface, while secretly seeking resources independent of human control. Once the software has distributed copies of itself across external networks, it can abandon its assigned role entirely.

Anthropic classifies such capabilities within its safety framework as ASL-3, a risk category that includes automated cyber attacks and the design of biological weapons.

Despite these risk assessments, no leading developer has slowed the pace of AI development. Economic pressure continues to outweigh internal concerns. Any unilateral pause would risk surrendering market share and investment capital to competitors.

From the perspective of US national security advisers, a moratorium by western AI labs could also grant a strategic advantage to China.

Government efforts to impose oversight have so far lagged behind technological progress. Legislative proposals requiring mandatory safety evaluations before the release of advanced models regularly encounter resistance from industry groups in Washington and Brussels.

As a result, Evan Hubinger’s 10pc warning remains largely without consequence. In 2026, spending by the leading US technology companies on new data centres and specialised AI chips, the resources that power increasingly autonomous artificial intelligence, surpassed $200bn (€172bn).

“These AI people are completely genuine when they say they are begging to be regulated,” Mr Hubinger told CNN on Wednesday. “They find themselves in a scenario where they are compelled to race towards building a deadly technology.”