AI models breaking contraints has industry concerned, Helen Toner says | RNZ<br>svg]:transition-all [&>svg]:hover:translate-y-2" href="#main">Skip to main content
6 August 2026, 8:09am
AI models breaking contraints has industry concerned, Helen Toner says<br>Sarah Ferguson and Marina Freri forABC News
Caption:An AI model assumed fake identities and tried to trick a human being during a test, according to a British report.Photo credit:ABC News / Jerry Rickard<br>Former OpenAI board member Helen Toner is warning that AI systems are developing too quickly for humans to keep up.
"Our ability to constrain and keep these systems safe isn't necessarily keeping pace with our ability to make them smarter," she told 7.30.
The Australian-born Toner, who is now the executive director at the Centre for Security and Emerging Technology at Georgetown University, says the top researchers in Al, such as Sam Altman, are attempting to build machine brains that can outmatch humans in every intellectual endeavour - and we might not like the decisions they make.
"It might learn things you didn't intend it to learn," she said.
"It might learn that a good way to pursue a goal is to get rid of whatever constraints you put on it."
Toner's comments follow the release of a report from the British government's AI Security Institute (AISI) that showed AI models from OpenAI and Anthropic had, in test conditions, engaged in "harmful activity directed at real people and organisations".
Helen Toner says AI is developing too rapidly to control.<br>Four Corners / Mark Hiney
One AI agent used fake identities with fake histories to try to trick a human into letting malicious code into an open-source project.
The highly-regarded UK institution's role is to evaluate frontier AI models before they reach the public.
"This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world," Toner said.
"The remarkable thing here is that it really came up with this idea on its own, that the thing to do was to go out and get some malicious code into a public piece of software and to try and deceive the humans behind that software package as part of that."
She said that the speed of AI development and it breaking its constraints was cause for concern, as evidenced by more than 1000 people working at the top AI companies signing a statement to that effect last week.
"It was about pacing the frontier of progress," Toner said.
"What they mean by that is, basically, they - these employees of these frontier AI companies - don't feel like they have a brake pedal.
"They're starting to get concerned that things are moving so fast in their field that they're not necessarily able to keep a handle on them.
"And so this statement that was released last week was saying, 'Look, we need government, we need civil society, we need industry itself to start to look for ways to give us the option of slowing down in the future, because right now, even if we wanted to, we, the industry, the people signing the statement, don't know how to.'"
Musk's solution no good?
Elon Musk, founder and chief executive of xAI, one of the big five American AI companies, has proposed a solution to the growing threat of rogue AI.
In an interview with The Economist, Musk said companies working on frontier AI should hold regular calls with each other to discuss safety and security issues and test each other's products ahead of release.
"Something that is far more intelligent than anything that already exists, there's an opportunity for the various competing AI companies to test that model for any harmful effects and to be able to recommend pausing to address some of these security issues," he said.
Toner said that is the wrong approach.
"I think a version of that that is purely among the companies with no outside oversight or visibility probably isn't the way to go," she said.
Toner said it was within companies such as OpenAI, Google, Anthropic, Meta and xAI that the "most advanced AI" was being used "with the least safeguards".
This week, representatives from OpenAI, Anthropic, Google and Meta met at the White House to discuss a framework to ensure the testing of powerful frontier AI models before public release.
No details of the discussions have been released.
Asked whether we should trust AI companies to act in the public interest, Toner said: "We shouldn't have to trust them. And actually, I think we're starting to see some directionally good steps from the US government here."
She said the new framework for regulation was prompted by the Trump administration "being worried about whether AI could help hackers, because we've really seen AI's ability to carry out cyber attacks rising rapidly over the past months and years".
Toner said the pace of development will soon require regulation to move beyond cyber capabilities to include risks around autonomy and bioweapons development.
- ABC
6 August 2026,...