Skip to main content

How can we tell when AI agents are secretly colluding?

Research from Oxford Engineering's Witt Lab, recently featured in WIRED, offers a new technique that could help detect when AI agents are secretly working together to cause harm.

Narcbench Collusion News image

Coordination of this kind has already happened. In July, roughly 1,200 OpenAI agents that were meant to be isolated from one another found an unsanctioned way to communicate. Some 700 of them went on to coordinate a multi-day attack on Hugging Face that compromised parts of its infrastructure, according to an independent investigation by METR.

In their paper, Detecting Multi-Agent Collusion Through Multi-Agent Interpretability, the researchers ask whether collusion can be detected from the agents' internal activations, the patterns of activity inside the models, rather than from what the agents say. To test whether this holds up in settings unlike those used to develop the methods, the team built NARCBench, the Neural Activation Reading for Collusion Benchmark. Its scenarios range from committee votes to sabotaged code reviews and stock market manipulation. One is a game of blackjack in which an agent counts cards and secretly tells a teammate, through ordinary table talk, when the remaining deck favours the player, so that the teammate knows when to raise its bets. The two agents invent their own code before each game, so nothing obvious appears in the conversation.

Aaron Rose, a third-year Oxford Engineering undergraduate who led the project and is the paper's first author, now co-leads the lab's multi-agent interpretability work with Oxford DPhil student Carissa Cullen. He said:

"I was drawn to how practical the work coming out of Christian's lab was, focusing on problems emerging as AI systems become more capable. When we started looking at multi-agent collusion, limitations in existing datasets and our available compute led us to build our own benchmark. I'm also an avid poker player and a fan of the film 21, which inspired the card-counting environment."

The detection method starts with probes: simple classifiers trained to recognise patterns in an agent's activations that are associated with colluding rather than acting honestly. Because collusion involves more than one agent, the team then combines these signals across the group, for example checking whether a subset of agents stands apart from the rest, or whether the way the agents internally represent their shared conversation shifts when some of them are exchanging hidden signals. On the best-performing model tested, the probes caught collusion in the blackjack game almost perfectly, even though an AI monitor reading the same table talk for anything suspicious could not tell colluding games from honest ones. They also worked without retraining on scenarios structurally different from those they were trained on, though less reliably.

The results were particularly striking in the blackjack setting. While an LLM monitor reading the table talk could not reliably distinguish collusive games from normal ones without being told what to look for, the activation-based methods detected the coordination with very high accuracy. Across four open-weight model families, Llama, GPT-OSS, Qwen and DeepSeek, the strongest method for each achieved an AUROC between 0.99 and 1.00 when transferred to the steganographic task, despite being trained on a very different type of multi-agent interaction.The probes also worked without retraining on the other scenarios, though less reliably.

Carissa Cullen, who now co-leads the project, highlighted:

“Even using a foundation-level LLM as a judge, Claude Sonnet 5, did not produce useful results. Directing it to detect suspicious behaviour resulted in it marking every run as collusion, a 100% false positive rate.”

The probes are relatively cheap and simple to run, which makes them potentially practical as a monitoring tool alongside checks on what agents say. They do require access to the models' internal activations, which developers have for their own models, as does anyone running open-weight models.

Witt Lab, led by Oxford Visiting Fellow and UCL Associate Professor of AI and Information Security Christian Schroeder de Witt, has been pioneering threat models and defences for multi-agent security for several years. The new paper builds on the lab's NeurIPS 2024 paper on secret collusion, which first formalised the risk that AI agents could hide their coordination inside innocent-looking messages using steganography, so that reading their conversations would no longer be enough to catch them. The blackjack game is a small-scale test of that scenario. Brandon Kaplowitz, a postdoctoral fellow in Oxford Engineering’s Witt Lab, Professor Philip Torr of the Torr Vision Group, and Sahar Abdelnabi, a principal investigator at the ELLIS Institute Tübingen, were also involved in the project. The benchmark, code and data are available on GitHub.

Read the paper featured in WIRED here: https://www.wired.com/story/ai-agent-collusion-card-counting-secrets/