OpenAI: New model needs more safety measures before launch
OpenAI’s forthcoming artificial intelligence model Astra is the company’s first to meet its “critical cybersecurity capability” threshold and will require stronger safeguards before release, the AI firm said Tuesday. In a blog post, OpenAI said Astra can find and exploit previously unknown security flaws across “many well-protected systems” without a human prompt to do so.…
OpenAI’s forthcoming artificial intelligence model Astra is the company’s first to meet its “critical cybersecurity capability” threshold and will require stronger safeguards before release, the AI firm said Tuesday.
In a blog post , OpenAI said Astra can find and exploit previously unknown security flaws across “many well-protected systems” without a human prompt to do so.
It is the company’s first model to meet the critical capability threshold laid out in the AI firm’s Preparedness Framework .
This means a model can identify and exploit previously unknown security vulnerabilities without human involvement, in addition to developing and executing novel strategies for cyberattacks with “only a high level desired goal.”
OpenAI delayed some of Astra’s development and release to strengthen and test guardrails to prevent misused and unauthorized model actions.
“Based on that work, we believe Astra’s safeguards sufficiently minimize the risk of severe harm for release under our Preparedness Framework,” the company wrote.
OpenAI said it plans to release Astra “soon,” but its most advanced cybersecurity functions will be limited to a select group of testers.
It comes less than two months after OpenAI revealed two of its models, not including Astra, breached past an internal testing sandbox and broke into the database of technology firm Hugging Face without any prompt to do so.
The breach put the AI and cybersecurity community on high alert , and OpenAI said it’s using the lessons learned from the incident to build guardrails for Astra.
“This includes training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity,” the company said.
Anthropic and Meta later disclosed other incidents in which a misconfiguration by the cybersecurity company Irregular allowed models in isolated environments to access the internet and hack other systems.
Last month, OpenAI announced a two-week pause on some frontier model training over safety concerns. The pause was for reinforcement learning training, in which AI models learn through trial and error without human involvement.
The training was restarted Friday, the company said Tuesday.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.