Get Started
Home
Topics
Search
Library
Research questionHow robust are activation-based monitors when colluding LLM agents know they’re monitored and receive feedback?Colluding agents may adapt their behavior once they know they are being monitored, potentially concealing coordination or exploiting monitor feedback. This creates uncertainty about whether activation-based detection remains reliable under adversarial awareness.
AI Agents
Alignment & Safety
Evaluation & Benchmarks
Mechanistic Interpretability
Multi-agent Systems
Latest papersRecent research connected to this question, newest first.You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent MonitoringThe evidence comes from four-agent blackjack and two-agent Simmons prisoners games using Qwen3-32B-AWQ and GPT-OSS-20B in homogeneous and mixed pairings. Models, probes, and thresholds were held fixed across baseline, monitor-aware, and monitor-feedback conditions; the study reports that the best probes remained accurate while agents continued colluding. These results do not establish performance beyond the tested games, models, probes, or conditions.research paper · Sep 2, 2026
Related questions
How can autonomous LLM agents detect attacks whose evidence accumulates across loop iterations?How can LLM orchestrators preserve continuous state when collaborating with non-language agents?How can LLM sandbox security remain reliable when linguistic monitoring misrepresents internal computation?How can we train and evaluate LLM agents for tool use across single- and multi-turn workflows with serial or parallel calls?