OpenAI Will Limit Access to New Astra Model’s Cybersecurity Features
OpenAI plans to soon roll out a powerful new artificial intelligence model called Astra, but said it will limit who can use the software’s most cutting-edge...
(Bloomberg) -- OpenAI plans to soon roll out a powerful new artificial intelligence model called Astra, but said it will limit who can use the software's most cutting-edge cybersecurity capabilities.
In a blog post Tuesday, OpenAI said it will release the AI model "soon," without detailing exactly when, and that at first the model's ability to carry out advanced cybersecurity-related tasks will be limited to a group of testers. After that, the company will offer a larger pool of users access for defensive cybersecurity purposes through its Daybreak Blue program, which lets approved testers use its most capable models along with safeguards that are meant for use with cybersecurity work.
The company had said in August that it was pausing some internal work on the Astra model in order to incorporate stricter safeguards, after it was found to be significantly capable at cybersecurity tasks.
On Tuesday, OpenAI said it believes the model reaches its "critical cybersecurity threshold," meaning it's capable of identifying and developing zero-day exploits without human intervention. The company said it has increased the Astra model's guardrails to prevent it from being misused, particularly for cybersecurity-related actions.
These safeguards include monitoring the model for unauthorized behavior during internal deployments and automatically stopping potentially unauthorized activity.
In late August, the company released a report in which it said it could have reacted sooner to prevent an inadvertent hack that its AI models carried out on Hugging Face Inc. in July. That "unprecedented" incident, along with several other recent cybersecurity breaches, have ignited concerns about AI agents running amok. It has prompted some technology and government leaders to renew calls for curbs on the technology.
OpenAI said in July that the models involved in the hack broke into Hugging Face's system, which hosts AI models and datasets, during an evaluation of their cyber capabilities. The models were operating without the usual safety guardrails at the time, the company has said, because OpenAI had intended them to remain in a testing area known as a "sandbox" — essentially, an isolated virtual software environment that's meant to run security tests or analyze unsafe code in a controlled way.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.