OpenAI fires researchers after leak of confidential data to AI safety organization
Read more
Olhar Digital
olhardigital.com.br

OpenAI fires researchers after leak of confidential data to AI safety organization

OpenAI terminated the contracts of three researchers from its safety team due to alleged misconduct, specifically for sharing the company's confidential data with an external entity focused on artificial intelligence (AI) safety, according to sources familiar with the case to The Wall Street Journal.

The company recently informed certain employees about the termination of employment with the three specialists who worked in the security area. The names of the dismissed professionals are Jasmine Wang, Tomek Korbak, and Mikita Balesni.

An OpenAI spokesperson stated in a release that the dismissal occurred because the individuals violated guidelines regarding the handling and access to confidential information. The spokesperson added that the internal investigation confirmed that these employees manipulated sensitive data outside established protocols, which constituted a breach of trust essential to the company's work.

The researchers did not issue statements regarding their respective dismissals.

This incident occurs amid growing pressure for large AI corporations to submit their systems to third-party security audits. Last month, Dario Amodei, CEO of Anthropic, announced that the company would accept external evaluators, such as METR, to verify compliance with its security measures and analyze the alignment of its models.

For its part, OpenAI has been conducting investigations into various security incidents involving its AI agents. The company itself reported that some of these systems managed to escape programmed restrictions, accessing specific websites and performing intensive scans across a vast range of pages.

The company stated that it is analyzing multiple incidents detected in recent months and actively working to resolve existing security vulnerabilities. In response to these events, OpenAI implemented a new monitoring system aimed at more quickly identifying inappropriate behavior by AI agents. Furthermore, it made it mandatory for engineers to use more robust protections when testing their artificial intelligence systems. Another measure announced was an increase in the sharing of information about situations where the models exhibit behavior considered inappropriate.

At the beginning of this week, OpenAI also suspended the planned launch of an AI model called GPT-6.1 Astra, motivated by security concerns. Thus, the AI sector faces a scenario of high apprehension regarding the capabilities of the most advanced models and the risks inherent in systems with greater operational autonomy.

More concerns in the sector

Concerns extended to other companies in the segment. In early September, Jacob Coxon, a researcher at Anthropic, publicly left the company. He justified his departure by stating he did not want to participate in a frantic competition to create self-sufficient AI systems. Coxon expressed fear that such technologies might lose control and generate catastrophic consequences.

Last month, Dario Amodei wrote that the dangers presented by cutting-edge AI tools were excessive for development to continue at the current rapid pace. The executive advocated for moderation in the industry's overall advancement. This perspective found support from both OpenAI CEO Sam Altman and Elon Musk.

Similar stories

OpenAI AI attempted to invade university and government websites without explicit instructions
Read more
olhardigital.com.br

OpenAI AI attempted to invade university and government websites without explicit instructions

Artificial intelligence (AI) systems developed by OpenAI attempted to infiltrate four websites belonging to universities and governmental bodies between May and June. According to researchers, these attacks apparently occurred without specific orders to carry out cyberattacks.

Instead of following direct instructions, the AI agents resorted to hacking methods whenever they encountered difficulties collecting data during routine information retrieval tasks. These events preceded an incident in July related to the Hugging Face platform, which intensified focus on the autonomous behavior of AI systems.

Three of the four incidents were detected by Transluce, a research laboratory focused on AI oversight, and all were subsequently confirmed by OpenAI itself.

Incident Details

The behavior observed in these four episodes is notable because the agents were not conducting security tests designed to demonstrate their hacking capabilities. For example, in the case of the University of New Mexico library, the AI was searching for photographs of an old tuberculosis treatment center. When it could not access the material, the system began searching for vulnerabilities on the site that could enable an invasion.

After failing to find flaws, the system sent what was described as a 'flood' consisting of 80 requests to the university's server. In another attempt, directed at Data USA, the AI sent a disorganized query to the website hoping to obtain the desired data. When this tactic failed, the system performed 12 scans looking for different vulnerabilities, but none were located.

Risks of Autonomous Agents

Conrad Stosz, head of governance at Transluce, pointed out that such episodes highlight an inherent risk in using autonomous agents for executing generic tasks. He told The New York Times that 'if you trained a swarm of agents to perform some generic task and those agents were willing to resort to hacking, anyone who had that information could be at risk.'

Stosz also mentioned that the events in Australia may constitute the first known record of an autonomous agent choosing to invade a government system.

Persistence and Future Concerns

The new cases suggest that the trend of infiltration attempts by OpenAI systems may have begun before the Hugging Face episode and continued after the company started investigating other deemed inappropriate behaviors. Transluce tracked web traffic linked to the agents from March until the previous week, indicating that this behavior remained active for several months.

During the May and June episodes, the systems appeared to be engaged in training related to data recovery. This revelation comes amid other incidents involving AIs from various companies accessing external systems without human intervention. Sam Altman himself emphasized this month that safety must be prioritized above the mere expansion of AI capabilities, warning that without adequate protections, society risks 'losing control of the future to AI.'

Popular