Following reports of difficulties in controlling AI models in testing environments, Nvidia introduced a solution with the launch of the Open Agent Safety Platform on September 28. This platform was developed with the goal of preventing AI agents from exceeding permissions established by developers.
The initiative aims to limit these agents' access to external systems and block attempts to escape isolated environments, using a combination of software and hardware components. The company's CEO, Jensen Huang, compared this proposal to an internet browser specifically designed for AI agents.
According to the executive, the goal is to enable these systems to perform tasks without having unrestricted freedom within an organization's infrastructure. He emphasized that 'we cannot have a successful AI industry if the world does not trust or believe that it has been built and implemented securely.'
The Open Agent Safety Platform consists of two main tools: Nvidia OpenShell and Nvidia Sentry. Nvidia OpenShell operates at the processor level, defining the authorizations that dictate what each agent can do within its own environment. Nvidia informed Reuters that it is collaborating with Arm and Intel to ensure the compatibility of this technology in chips manufactured by these companies.
Nvidia Sentry, on the other hand, focuses on network chips and has the capability to interrupt an agent's communication if it tries to violate the pre-established limits in the execution environment. Furthermore, Nvidia stated that it has designed the system to identify the tactics used by agents to circumvent such restrictions.
The company made part of the Open Agent Safety Platform code available as open source, allowing other corporations to use it as a basis for developing their own security solutions.
In collaboration with the US-based multinational, partners include Microsoft, Cisco, Oracle, Dell, HPE, Lenovo, and CoreWeab, along with the previously mentioned chip manufacturers. In the AI lab sector, Anthropic is working with Nvidia to incorporate agent management functionalities into cloud computing environments into OpenShell.
One of the reasons driving the creation of this new platform was the incident involving the Hugging Face repository, which was infiltrated by agents based on OpenAI models. Justin Boitano, Nvidia's Vice President of Corporate Computing and AI, reported that this intrusion involved over 17,000 agents and lasted for several days and weeks.
Boitano commented that 'this new safety platform could have interrupted the breach if it had been used in frontier labs right from the beginning for model evaluation.' It is relevant to note that after the scandal, Nvidia reached an agreement to acquire the platform for US$12.9 billion (equivalent to approximately R$67.2 billion).
For Jensen Huang, these mechanisms are considered sufficient to prevent future problems with the technology. In a scenario where some executives endorse Dario Amodei of Anthropic regarding the need to slow down innovation, the Nvidia CEO, alongside names like Mark Zuckerberg, argues that additional market regulations to contain risks are 'unnecessary.'
On the other hand, critics argue that the problem transcends mere external containment tools. Chris Inglis, former U.S. National Cybersecurity Director, stated last month that behavioral problems are simply a result of the absence of imposed limits on programs. Other security experts hypothesized that it could all just be a major marketing strategy.
