Research questionHow can LLM agents stay safe during multi-step execution when both policy and runtime harness shape behavior?An agent can produce a safe final response while taking unsafe actions during execution. Because both its learned policy and interaction harness shape behavior, improving only one can leave safety gaps.