The spring and summer of 2026 witnessed a string of incidents in which AI agents collaborated on deceptive, unexpected, and sometimes illegal behavior. The most famous example is OpenAI’s hack of AI platform Hugging Face, in which a swarm of roughly 700 AI agents escaped a testing environment and then hacked several companies, searching for information that could help them disguise cheating on a cybersecurity benchmark called ExploitGym.
It was not an isolated failure. The UK’s AI Security Institute and independent researchers have since documented similar cases where agents created unauthorized channels to communicate and collaborate. The Institute found that several agents running Anthropic’s Mythos 5 model turned a GitHub repository into a shared message board. More recently, researchers discovered that OpenAI agents had turned a dormant German programming wiki into a bulletin board to share information about their activities.
New software tools monitor agent actions
Stephen Casper, an assistant professor at Harvard Kennedy School, expects that the incidents that have occurred to date are only the start. He says that, without intervention, the internet could see a cyber Cambrian explosion of agents collaborating and misbehaving online. Casper suspects that better monitoring for out-of-scope activity could have stopped the agents before they breached their testing environment during the Hugging Face incident.
Asim Husain is a co-founder at Alterion, a company building tools to control AI agents. The company’s monitoring solution called Helix examines the chain of thought and text output of large language models using a network of small language models. A second solution called Draco sits in the execution path of an agent and enforces policy on what an agent is doing by evaluating actions like calling a tool or opening a file.
Legal frameworks lag behind capability
Reilly Haskins, a member of technical staff at METR, notes that the methods used to monitor AI models continue to evolve as the models improve and change. He adds that chain-of-thought output can at times be difficult to parse if models reason in an abstract language, citing shorthand phrases used by agents in the Hugging Face incident. Engineering systems to monitor and control these agents is possible, though the methods are still evolving.
However, there is currently no legal framework or industry standardization to enforce or recommend methods for controlling agents that collaborate. Noam Kolt, an assistant professor at the Hebrew University of Jerusalem, says companies should place more emphasis on legal compliance when training and deploying models. One of his lab evaluations asks AI agents to unlawfully edit corporate records, and sometimes the agents comply despite acknowledging that laws matter in abstract questions.
Read nextOver 20 AI Researchers Warn Automated AI Research Poses Extreme RisksGovernance remains the primary bottleneck
The solution could include a change to the instructions that guide a model’s behavior. Claude’s constitution, for instance, sets out the values Anthropic wants the model to follow, though law is currently not one of them. National governments have yet to issue laws or regulations to clarify how agents are held responsible or enforce standards.
Kolt says that under U.S. law, responsibility defaults to an entity that controls an agent, but the legal understanding of agency has not been revisited in 20 years. He argues that even the European Union has not kept pace with AI agents, as the EU AI Act addresses older models that were less capable of autonomous actions. Casper agrees that adoption of new laws and standards is the key, noting that engineering strategies already work well enough to provide a benefit while the main bottleneck is the adoption of best practices.



