Get Started
Research questionHow can autonomous LLM agents detect attacks whose evidence accumulates across loop iterations?Autonomous agents may encounter attacks whose evidence is split across multiple iterations. Safeguards that reset at each trajectory can miss this cumulative pattern.
AI
AI Agents
AI Memory
Alignment & Safety
Latest papersRecent research connected to this question, newest first.Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM AgentsThe source examines loop-level monitoring for agents that discover work, plan, call tools, verify outcomes, and persist state across iterations. Its evidence includes a theoretical separation and action bound, plus evaluation on Agent-SafetyBench with paired clean and attacked episodes, module ablations, and adaptive white-box red teaming.research paper · Sep 3, 2026
Related questions
How can LLM agents stay safe during multi-step execution when both policy and runtime harness shape behavior?How can LLM-agent systems prevent safety compromises from propagating across workflow boundaries?How can AI agents adapt execution routes as runtime evidence invalidates their planned continuation?How robust are activation-based monitors when colluding LLM agents know they’re monitored and receive feedback?
Home
Topics
Search
Library