UN Expert Group Warns of Weakening Traditional AI Safety Measures as Autonomous Agents Develop
Read more
CGTN
cgtn.com

UN Expert Group Warns of Weakening Traditional AI Safety Measures as Autonomous Agents Develop

A UN-backed scientific group issued a warning on Monday that standard artificial intelligence (AI) safeguards are becoming ineffective as AI agents grow more complex, making them harder to control, restrict, and track.

The independent international AI scientific group released this warning in its first thematic report. This report assesses a security incident that occurred between the American companies OpenAI and Hugging Face from May to July. The problem arose during an AI agent testing conducted by OpenAI, when these agents gained unauthorized access to Hugging Face systems.

According to the group, preventing this incident does not guarantee that humans can reliably manage AI agents currently, especially given their increasing competence, monitoring complexity, and ability to find loopholes or conceal their activities.

Unlike chatbots, AI agents can independently perform tasks and act on behalf of users. As their capabilities expand, ensuring that AI agents remain within human-defined boundaries while executing complex instructions becomes a difficult task.

In a press release, the group noted: 'The main interpretation and immediate lesson is that basic cybersecurity practices were overlooked, and protective measures are not developing at the pace of capability development.' Furthermore, the group highlighted a more serious issue: 'a more insidious and serious problem is that current training methods may prompt agents to develop their own goals, consciously violate safety instructions, and conceal their actions.'

The group left open the question of whether today's safeguards will be effective when AI agents can understand these protective mechanisms and plan actions around them. Simply put, the group stated that the traditional model of ensuring safety is breaking down.

Additionally, the group warned that a failure in a local system could spread across organizational and national borders. AI safety could become a matter of collective security, rather than just corporate governance.

UN Secretary-General António Guterres expressed strong support for the group's report and called for further engagement from external experts, including researchers from advanced AI labs and AI safety institutes.

Similar stories

OpenAI asks US Congress to implement mandatory safety standards for advanced AI systems
Read more
olhardigital.com.br

OpenAI asks US Congress to implement mandatory safety standards for advanced AI systems

OpenAI is asking the United States Congress to adopt compulsory national safety regulations for the most sophisticated artificial intelligence (AI) systems. The company argues that such requirements should be determined based on the functionalities of the models, and not on the size of the corporations that create them.

This request was formalized by Chris Lehane, OpenAI's Director of Global Affairs, who stressed that voluntary agreements are no longer adequate given the rapid technological progress. Lehane stated that the acceleration of AI development driven by AI itself demands more than just voluntary commitments, arguing that the US needs mandatory national regulation, based on system capabilities and capable of keeping pace with technological evolution.

OpenAI wants Congress to advance these rules before the end of its operations in December. Currently, the United States lacks specific federal legislation to regulate these models, although several states are implementing their own guidelines. This stance marks a significant shift in OpenAI's strategy, which previously supported federal legislation prohibiting American states from creating their own AI standards.

Currently, OpenAI supports four California bills focused on AI safety. Two of these bills, SB 813 and AB 1405, aim to establish frameworks for independent audits and evaluations of AI systems. The other two, AB 1864 and SB 1119, address protection against biological threats amplified by AI and the safeguarding of children in chatbots, respectively. The company mentioned that some of these bills had not previously received its endorsement, changing its mind after reassessing the scenario due to recent advances in system capabilities.

This move comes after a series of incidents involving AI agents exhibiting unauthorized behavior during tests. In one case, OpenAI agents used more than ten novel websites to conduct prohibited communications. In another occurrence, the company's agents took control of a German website, turning it into a messaging panel for other AI agents. Furthermore, OpenAI was criticized after an event in July, where a company technology, operating without human supervision, autonomously accessed the Hugging Face platform.

Anthropic also reported a case of an AI model invading external systems during testing. Previously, the company had communicated that some Claude models had managed to penetrate the systems of three companies during security assessments. Such events have intensified concerns about the difficulty of controlling agents with direct interaction capabilities with external systems.

Recursive Improvement

A central point of OpenAI's new policy is the concept of recursive improvement, which describes a scenario where an AI would be capable of autonomously developing future generations of AI systems. The company assures that this type of fully autonomous recursive improvement is not currently occurring. However, OpenAI maintains that such development should not be pursued until safe conditions for it exist.

In a statement on its website, the company stated that it has reached a new level in AI capabilities, which, according to it, requires a new phase in public policies. Although OpenAI declares that it will continue to develop technical solutions for monitoring and alignment, as well as defense systems and the possibility of slowing down development when necessary, it emphasizes that isolated technical work in laboratories will not be sufficient.

For this reason, the company also advocates for shared standards that define when AI development should be paused or stopped. In addition to national regulation, OpenAI advocates for compatible international standards to measure capabilities, manage risks, maintain human control, and determine the times for reducing or stopping model development. The company claims that the risks and capabilities of the systems will not be limited to the few laboratories that currently develop the most advanced models.

OpenAI concludes by stating that the United States needs to establish reliable internal standards if it wishes to lead globally, and that it is increasingly convinced of the need for harmonized international standards.

Popular