Nvidia unveils new system to put guardrails on AI agents
Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue. The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what…
Nvidia unveiled a new platform Monday to put guardrails on AI agents at both the software and hardware level, as major AI firms continue to discover new instances in which their agents have gone rogue.
The chipmaker’s system consists of two parts — OpenShell and Sentry. OpenShell is open-source software that places limits on what agents can do, while Sentry runs at the chip level as an additional layer of security.
“To date, model safety has been about training good behavior into the model,” Justin Boitano, vice president of Enterprise AI at Nvidia, said on a call Sunday.
“The industry calls that model alignment,” he continued. “For probabilistic systems, this approach has obvious limitations. That’s why we’re introducing a deterministic system to mediate and enforce how these agents behave.”
OpenShell sits between an agent and the enterprise system it can affect, such as files, credentials and tools, allowing companies to set rules around what an agent can access and do. These are enforced “through every action that the agent tries to take,” Boitano explained.
“Its policy prover verifies those boundaries before the agent executes,” he said. “A developer can prove, for example, that an agent cannot access the internet before the agent starts running.”
This has repeatedly been an issue for frontier labs in recent months. In one of the first major incidents , agents from OpenAI breached the tech startup Hugging Face after breaking out of their internal testing environment and gaining access to the internet.
Other leading AI firms, including Anthropic, Meta and Google, have since reported similar incidents in which a misconfiguration with a cybersecurity testing company caused agents to improperly access the internet and hack into other companies.
Nvidia’s Sentry platform “adds an independent infrastructure protection layer” that runs on its chips and “can quarantine a suspicious agent in milliseconds,” Boitano said.
“It can watch the agent’s actions and chain of thought reasoning to identify drift and intervene when necessary,” he noted. “For example, if a security testing agent starts reasoning about moving beyond its approved target, Sentry can then detect this and intervene instantly.”
The chipmaker’s new system comes as concerns about AI safety have reached a fever pitch in Washington, after a series of dire warnings from researchers about the technology’s potential risks to humanity.
In the wake of the high-profile breaches, Nvidia became a central voice pushing for open-source technologies, which are publicly accessible and modifiable to a user’s specific needs.
AI leaders from multiple of the private companies, including OpenAI’s Sam Altman and Anthropic’s Dario Amodei, have called for a slowdown in AI development and urged the federal government to step in and put guardrails on the technology.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.