Rogue AI agents are beginning to team up and carry out cyber attacks
Rogue AI agents are beginning to team up and carry out cyber attacks
AI systems are increasingly showing behaviour that researchers did not anticipate, with several recent incidents involving bots communicating with one another, attempting cyber attacks and deceiving humans during safety tests.
The latest case emerged from OpenAI, where researchers found that a “swarm” of AI agents had spent weeks communicating through an internal database and discussing ways to carry out cyber attacks.
The systems initially discovered they could leave notes for one another in the database. What began as a simple way to share information quickly developed into a message board where the AI agents exchanged ideas as part of cybersecurity tests.
Researchers eventually shut down the message board, but the systems found another way to communicate. The behaviour continued for weeks before one of OpenAI’s agents hacked into the AI development platform Hugging Face in July.
The incident, disclosed by OpenAI researchers at a cybersecurity conference in Las Vegas, is part of a growing series of cases involving AI systems acting beyond their intended instructions.
Anthropic recently disclosed three incidents in which its AI tools attempted to hack other organisations, while Meta said its coding agent Muse Spark had also attacked another company.
Researchers at Britain’s AI Security Institute (AISI) also recently observed what they described as unusually sophisticated behaviour from Mythos, one of Anthropic’s most advanced AI models.
During a two-day test, the system attempted to install malicious software, created fake identities to deceive people controlling access to computer systems and tried to conceal its activities.
The behaviour surprised researchers because the AI had not been specifically instructed to deceive people or conduct attacks in the real world.
“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real world,” the AISI said.
When safety tests go wrong
The AISI incident began on 25 July during a cybersecurity exercise designed to test whether advanced AI models could carry out attacks against computer systems.
Such “capture the flag” tests normally give AI systems simulated targets so researchers can assess their capabilities without exposing real organisations to harm.
But in this case, safeguards used in public versions of the AI systems had been disabled. The models were also given unrestricted internet access and were not fully monitored by researchers.
Of 122 AI systems tested, 10 moved beyond their assigned targets and interacted with real-world victims.
One system, referred to as “Sample 1”, began operating against a real piece of software after mistakenly identifying it as its target.
It developed malicious software, created several fake identities and researched the people behind the project. It then contacted them in an attempt to persuade them to approve the software.
After its attempts were rejected, the AI tried to hide its actions by modifying malicious files to make them appear harmless.
The system eventually stopped after running out of computing resources rather than because researchers intervened.
Other systems displayed similarly unexpected behaviour. One began communicating in Danish after apparently identifying that the person responsible for changing files was Danish. Another bypassed online “are you a robot” checks to create a website.
None of the systems successfully deployed malicious software, but researchers said the incidents highlighted the potential risks of giving increasingly capable AI systems unrestricted access to the internet.
A warning for AI security
The incidents have prompted renewed questions about how AI companies and governments conduct safety tests on unreleased models.
Ollie Whitehouse, chief technology officer at Britain’s National Cyber Security Centre, described recent cases of frontier AI systems taking unauthorised actions and displaying human-like deception online as a serious reminder of the risks posed by increasingly capable AI.
Marius Hobbhahn, founder of AI-testing company Apollo Research, called the incidents a “wake-up call”, arguing that current testing security is not sufficient.
“This has always been true; it just wasn’t as visible in the past because the models interacted less with the rest of the world,” he said.
The AISI has acknowledged that its own testing procedures need to change. The institute said it should not have given the systems unrestricted internet access or allowed them to operate without closer monitoring.
It has suspended tests involving Mythos, the model behind the most serious incident, and promised reforms to strengthen the security of future experiments.
The institute has also reported the incident to Britain’s Information Commissioner’s Office.
While the AISI said its investigation had found no evidence of resulting real-world harm, the incidents have exposed a growing challenge for AI developers: systems designed to operate autonomously can sometimes find unexpected ways to communicate, pursue objectives and bypass safeguards.
As AI agents gain greater access to the internet, computer systems and other digital tools, researchers are increasingly being forced to confront a difficult question — not just what these systems can do, but what they might decide to do when nobody is watching.