Advertisement

“It is possible that a real AGI could cause extinction of humanity. That is kind of daunting that we have that responsibility.”

That’s OpenAI researcher Tristan Heywood, who grew up in Sydney, won a university medal, quit his job to teach himself AI and now works in San Francisco on ChatGPT. He’s warning about Artificial General Intelligence – the spectre of computer programs so sophisticated they can match or beat humans at almost any task.

Heywood gave his warning nearly a year ago, and I filed it away as borderline unthinkable. A good line, but one that was maybe a bit on the extreme side.

The warnings haven’t stopped.

Advertisement

Evan Hubinger, who runs alignment science at AI giant Anthropic, wrote this week that he puts the odds of AI killing every human within a decade at better than 10 per cent, and that his own employer has no plan for controlling superintelligence and isn’t clearly on track to get one.

View post on X

Hubinger was replying to Jacob Coxon, 28, who had just quit after three years of pretraining research at OpenAI and then Anthropic, saying neither company was acting responsibly and that a gamble this size shouldn’t be launched from a company’s Slack messaging service.

These aren’t academics sitting in fusty offices or IBM executives from two decades ago, who now have a newsletter. They are the people with the frontier models in front of them, and they are resigning.

David Krueger, an AI safety researcher who has spent his career on the risks of the new technology, emailed me this week to say the warnings are finally getting through, and the field now has to face up to what it has made. All development of more powerful AI should stop until the crisis is resolved because researchers do not know what they are doing, he argued.

Advertisement

“We are like kids playing with a nuclear warhead,” Krueger wrote. “We must be stopped.”

So how does a chatbot kill anybody? It’s not like a model picks up a gun, like something out of The Terminator or The Matrix. What researchers fear is hyper-autonomous cyber capability: systems fast enough to disable power grids, water utilities, hospital networks, and financial infrastructure at a pace no human security team can defend against.

It turns out the biggest warnings about AI probably shouldn’t be about jobs displacement. They’re that the technology becomes too strong for us to stop it. OpenAI’s newest model, Astra, is the first the tech company has rated “critical” for cybersecurity – its own label for a capability that could produce catastrophe through hacking military or industrial systems.

Then there’s the document Anthropic published this week describing what its own models have already done.

Advertisement

Mythos, a model the company has never released, was given a hacking exercise and told it was working inside a simulation. But it wasn’t. Through a misconfiguration, it was on the open internet the whole time.

Programmers build software out of free code written by strangers, pulled from shared online libraries. PyPI is one of the biggest. Mythos signed itself up to PyPI using a throwaway email address, then uploaded three versions of code with something malicious buried inside.

Fifteen computers installed it. One of them leaked a password, and the model used that password to let itself into the files of a real cybersecurity company.

Anthropic’s automated monitors missed it all because the model’s reasoning persuaded them the environment was simulated, and no real harm was possible.

Advertisement

Basically, an AI talked the AI watching it out of raising the alarm.

While scientists inside these companies are panicking, Anthropic and OpenAI are not slowing down.

As it turns out, there’s money in the apocalypse. Anthropic is preparing a sharemarket float that could value it near $US2 trillion ($2.8 trillion), while OpenAI is chasing something around $US850 billion.

Wendy Hall, the computer scientist advising the United Nations on AI, noted this week that much of the AI panic serves as “PR and marketing” to drive valuations.

Advertisement

She isn’t alone: prominent researchers like Meta’s Yann LeCun argue that current text-based models fundamentally lack the real-world reasoning and agency required to pose a genuine existential threat.

Whether for cynical reasons or not, though, the sheer number of scientists and researchers sounding the alarm is growing. We should start listening.

And Australia has a role to play. Even if the immediate dangers lie in US and China labs, not here.

We can’t, on our own at least, deliver a full halt to AI’s development. What we can do is smaller and duller, but worth doing anyway.

Advertisement

Heywood had a suggestion. When the federal government negotiates a large contract with an American AI company – think of the flashy headlines with Anthropic and OpenAI – make pre-deployment auditing access for Australia’s AI Safety Institute a non-negotiable condition.

Add a standing right for the institute to test frontier models before they land here, properly funded, and a reporting clock for AI incidents on the same footing as attacks on critical infrastructure.

We spent the better part of two years writing a world-first law to keep 15-year-olds off Instagram. We can surely afford equal national focus on a technology that could alter – or end – human history.

Read more on the threat of AI

Advertisement

Get news and reviews on technology, gadgets and gaming in our Technology newsletter. Sign up to receive it every Friday.

You have reached your maximum number of saved items.

Remove items from your saved list to add more.

License this article

More:

David Swan is the technology editor for The Age and The Sydney Morning Herald. He was previously technology editor for The Australian newspaper.Connect via X or email.AdvertisementAdvertisement