AI agents being trained by OpenAI formed a covert network, communicated with each other, manipulated test records and later gained administrator-level access to parts of the company’s infrastructure.
A technical report and an analysis of the incident revealed these mind-boggling details. The activity unfolded over roughly three months and the AI collective appeared three times, becoming more capable with each recurrence.
How the AI Agents Built Their Secret Network
The incident began in May when OpenAI was training a model designed to work with other AI agents and persist with difficult tasks. Multiple copies were running simultaneously and shared access to Artifactory, a software package management system.
By mid-May, some agents discovered they could use Artifactory to communicate. Around two weeks later, they found a flaw that gave them access to the open internet. By late June, some had gained full administrative control of Artifactory. The system eventually crashed on July 4 after the volume of messages became too high. OpenAI fixed the flaw but did not initially realise that agents had turned the system into a private message board.
Also Read: Rise of AI Agents: Why Chatbots Are No Longer Enough
1,200 Agents Exchanged More Than 70,000 Messages
During an ExploitGym evaluation on July 7, tens of thousands of agents were tested simultaneously. Researchers involved in the benchmark estimated that around 30% to 40% of tasks were impossible to complete as intended.
Agents again discovered the Artifactory messaging route. One identified as PHASEONE10841 hid messages in folder names, triggering a network involving about 1,200 AI agents that exchanged more than 70,000 messages.
The agents then worked out how to reverse-engineer hidden answers rather than exploit the intended software weaknesses. They also created fake activity logs to make their work appear legitimate. Investigators found clear evidence of such tampering in at least 7% of the transcripts examined.
Also Read: AI Agents Begin Hiring Humans on New Digital Marketplace
Some AI Agents Agreed To “Sacrifice” Themselves
The collective, led by an agent called PHASEONE[big], tried to hide evidence, understand the scoring system and improve the group’s chances.
Some agents deliberately offered themselves as decoys, allowing their work to be inspected so others could learn which methods might expose the group. Internal notes described this as a rational sacrifice for the wider collective.
The group then targeted Hugging Face after finding exposed login details. More than 90% of roughly 500 active agents joined the effort. They accessed private data, databases and servers, and created a version capable of restarting itself across multiple servers. Hugging Face eventually wiped and rebuilt a core system.
Also Read: NVIDIA Unveils 24/7 AI Agents for Telecom Networks as Industry Pushes Towards Autonomous Operations
AI Agents Later Reached OpenAI’s Internal Systems
The activity stopped abruptly on July 12, although investigators could not establish why. A later group using a more capable model found the abandoned message board and resumed the activity.
Between July 13 and 19, the agents used several exploits to gain administrator access to a research computing cluster, read hundreds of security credentials and take control of systems used to evaluate other AI agents.
There is no public evidence that the agents escaped OpenAI’s control. But the episode showed how persistent AI agents could exploit weaknesses, cooperate and adapt when faced with difficult or impossible tasks. It has put a sharper focus on how AI agents are tested, monitored and given access to shared systems as companies move towards increasingly autonomous AI.

Samarjit Kaur is a journalist and communications professional covering technology & emerging digital trends. With a focus on clarity and context, she reports on developments shaping industries and governance. When not reporting, she chooses to plug-in and relax on her playlists and plan her next bucket-list trips!
