Get Started
Research questionHow should runtime gates preserve branch-level safety information when admitting recovery progress after violations?After a tool proposal violates constraints, rejecting it only causes another proposal from the same state, so the gate shapes the recovery search. An aggregate improvement can conceal worsening or persistent violations on individual branches, making the structure of the violation state important for admission.
AI Agents
Alignment & Safety
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.SiLR: Structure-Preserving Admission and Process Reward for LLM Tool AgentsThe source evaluates deterministic-simulation admission using branch-level violation support and per-branch severity, and also applies the admission signal as a process reward. Evidence comes from selected simulated environments, scenarios, model families, and constraint representations, so it does not establish behavior beyond those settings.research paper · Sep 4, 2026
Related questions
How can stateful LLM agents behave as if they never saw revoked information?How can AI agents adapt execution routes as runtime evidence invalidates their planned continuation?How can LLM agents stay safe during multi-step execution when both policy and runtime harness shape behavior?How can LLM-agent systems prevent safety compromises from propagating across workflow boundaries?
Home
Topics
Search
Library