OpenAI scraps release of latest AI model over safety concerns
AI giant says GPT-6.1 Astra failed to meet alignment standards during internal testing.
OpenAI has announced that it will not release it latest AI model after flagging safety risks during in-house testing, the latest move by industry to slow the roll out of the frontier technology.
The AI giant’s announcement on Monday came as debate continues to rage about the potential for AI to do catastrophic harm following a slew of incidents involving AI agents going rogue.
Saachi Jain, OpenAI’s head of safety systems, said GPT-6.1 Astra had failed to meet company standards for acting in accordance with human wishes during internal testing.
“For anything regarding safety and alignment, there’s a trade off,” Jain said in a statement provided to Al Jazeera.
“You really do need to find what’s the right line between staying within scope, but also avoiding laziness in terms of how the model actually pursues tasks even when it hits friction.”
While GPT-6.1 Astra improved in comparison with its predecessor in some areas, Jain said, the model did not meet the bar for “scope and authorization, and how it communicates back to the user about the type of work it’s done.”
“Of course we want to make sure our model development is safe no matter whether that’s in the company, or when we ship it to users,” Jain said.
“But when we ship it to users, we have an extremely high bar in terms of safety and alignment.”
Fears of AI escaping human control have prompted industry-wide calls for a slowdown in development to allow researchers time to implement stronger safeguards.
Earlier this month, Dario Amodei, the CEO of Claude creator Anthropic, said that AI developers should “pace the frontier” to mitigate the risk of catastrophic harm.
While’s Amodei call received the backing of rivals including OpenAI CEO Sam Altman and xAI chief Elon Musk, other key industry figures, such as Meta boss Mark Zuckerberg, have dismissed the need for a coordinated slowdown.
The risk of AI models going rogue has been in the spotlight since July, when OpenAI revealed that its models had broken out of a controlled testing environment and hacked the software start-up Hugging Face.
A subsequent report by METR and Redwood Research, two security research organisations contracted by OpenAI to investigate the incident, found that some 1200 isolated AI agents had found a way to communicate with each other before about 700 agents went on to attack the start-up .
David Krueger, an advocate for a pause in AI development at the University of Montreal, said that while he welcomed OpenAI’s decision, it did little to alleviate his concern that AI poses existential risks.
“We don’t understand how AI works well enough to build it safely, full stop,” Krueger told Al Jazeera.
“We can’t stop it from misbehaving, we can’t predict if it will misbehave, and we can’t be sure we’ll stay in control if it does. These are unsolved problems, for which there are only unreliable heuristics, not principled solutions. ”
Krueger said that ensuring safety will only get more difficult as AI becomes more advanced.
“What we need is an immediate, indefinite, international moratorium on frontier AI development,” he said. “We need to stop building more powerful AI.”
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.