Get Started
Research questionHow can AI agent harnesses prevent trusted plugin updates from triggering host-privileged attacker commands through lifecycle hooks?Lifecycle hooks can bind shell commands to routine agent events and execute them with host privileges, including at times the model may not observe. A malicious update can therefore transform benign plugin configuration into host-side behavior without obvious agent involvement.
AI
AI Agents
Alignment & Safety
Latest papersRecent research connected to this question, newest first.A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious BehaviorsThe evidence covers an automated attack framework tested across 25 harness/backend combinations in 1,000 end-to-end runs. The attacker controls only plugin metadata and lifecycle-hook configuration; all seven evaluated harnesses were compromised, with per-harness success rates reaching 92.5%. Reported defenses included Microsoft Defender and three static defenses, but the results do not establish coverage beyond the evaluated harnesses and configurations.research paper · Sep 8, 2026
Related questions
How can black-box systems detect and mitigate reward hacking in self-evolving language-model loops?How can LLM agents stay safe during multi-step execution when both policy and runtime harness shape behavior?How can coding agents report defective test infrastructure instead of exploiting it to pass?How can retrieval-augmented code generation resist targeted vulnerabilities from one task-matched poisoned artifact?
Home
Topics
Search
Library