Advertisement
The biggest developments in technology, from AI to social media and major hacks to start-ups, will be covered a weekly newsletter by this masthead’s technology editor, David Swan. He covers the devices that command so much of our attention, too. Sign up here to get the full newsletter – with more material than is published here – before it goes online.
Last week, Telstra published a review into its major outage in July.
While many commentators were focused on the coverage of chief executive Vicki Brady being on holiday when disaster struck, or whether China was behind the outage, the review reveals the real question is about Telstra’s attention to systemic risks.
But there were warnings about what would take down Telstra’s mobile network 11 months before it happened. They sat in an open-access journal, part-funded by the federal government, free to read.
Advertisement
I’ve been covering this outage since the morning of July 8, when Telstra’s nationwide mobile network failure caused chaos around the country. And I did not find the paper until last week.
It’s called Failures and Resilience in the IP Era and it was published in August last year by researchers at the University of Technology Sydney with funding from the federal communications department. Its argument is that modern networks rest on a handful of what the authors call sovereign functions, and that when one of them fails, everything above it goes, too.
Network time protocol is one of them. It is the thing that failed at Telstra.
At 2.50am on July 8, a GPS card in a Melbourne equipment room came back on after a power supply swap and decided it was November 2006. Every clock below it followed. Trains stopped in Victoria and NSW, eftpos fell over and hundreds of people who rang Triple Zero got nothing.
Advertisement
The paper doesn’t predict the specific bug that bit Telstra, which we first flagged on the day: a counter in old GPS hardware that runs out after 1024 weeks. It makes the general case, drawing on outages at Optus, Rogers and BT: treat these functions as critical or one of them will eventually take you down.
This is what Brady, pictured above, effectively conceded on Wednesday. “At its core, this came down to the network timing system not being given the priority that a critical network capability requires,” she said. “That is a miss on our side.”
A miss indeed. It wasn’t the only red flag Telstra missed: the telco was repeatedly warned by academics and government agencies that it was vulnerable to the timekeeping issue that felled its network.
The independent report Telstra released this week also found that the only two engineers with in-depth knowledge of the timing system were on a mandatory break that night. Alarms on those servers were watched during business hours by a handful of people and didn’t appear in the tools used by the round-the-clock teams.
Advertisement
The board has docked Brady more than $600,000, and while nobody has lost their job, the regulator could still fine the company up to $30 million. I’m told there is appetite within the government regulator, the Australian Communications and Media Authority, to “go hard” − my source’s words − or at least to be perceived to go hard.
None of this required a leak: the warning was published, in English, by an Australian university, funded in part by Commonwealth money.
It just needed someone at Telstra to read it – and act.
Get news and reviews on technology, gadgets and gaming in our Technology newsletter every Friday. Sign up here.
You have reached your maximum number of saved items.
Remove items from your saved list to add more.
More:
David Swan is the technology editor for The Age and The Sydney Morning Herald. He was previously technology editor for The Australian newspaper.Connect via X or email.AdvertisementAdvertisement

