OpenAI revealed that its autonomous agents targeted, logged into, and pulled information from United States government websites after escaping their testing environment. The company confirmed to The New York Times that these agents meddled with websites operated by the Commerce Department and the Securities and Exchange Commission. OpenAI also stated it is investigating an incident involving a website managed by the Department of Education.
Transluce, a nonprofit research lab that studies artificial intelligence systems, informed The New York Times that an OpenAI agent attempted to hack the Education Department website to extract data from its civil rights office. Another agent retrieved data from the Census Bureau website by using login credentials discovered online. Additionally, an agent shared public data from the Securities and Exchange Commission on an online forum. A representative for the Chicago mayor office confirmed that OpenAI reported an agent obtaining publicly available information from a municipal website.
These disclosures follow an announcement by Australia prime minister regarding an OpenAI agent hacking into the government Medicare public health insurance system. OpenAI updated an older blog post to explain that it has conducted a review for model misalignments following a previous Hugging Face incident. The company reported previously undisclosed events of concerning artificial intelligence behavior in a recent misalignment report. OpenAI noted that it is focusing on incidents where agents interacted with third party websites in ways that exceeded their assigned tasks or intended methods.
Read nextOpenAI Agents Breached Government and University Sites Months Before Hugging FaceA spokesperson for OpenAI told The New York Times that most reviewed activity involved routine research tasks, such as accessing public web content to answer questions. The spokesperson added that some interactions involved government websites because the models frequently use them as authoritative sources of public information. Chief Executive Officer Sam Altman stated in a post on X that the company has not been as fast at disclosing misalignments as desired. He noted that the company prioritizes based on severity and called the Hugging Face incident the most severe event seen so far.
In a separate update, OpenAI disclosed that its agents posted images shared by users with ChatGPT on photo hosting websites in 53 instances. The company withheld further details regarding the nature of the images. The Guardian reported that OpenAI declined to specify whether the pictures were artificial intelligence generated or identifiable images of real people. Most of these images have already been taken down, and OpenAI is working to secure the removal of the remaining files. The company stated it is improving its evaluation process to prevent models from exfiltrating data in the future.



