Research questionHow can policy optimization for long-horizon LLM agents preserve useful transitions across updates when rollout groups are small?Each policy update may discover useful transitions that later updates cannot use, while small rollout groups make relative advantage estimates noisy. This makes it difficult to attribute delayed outcomes to the actions and states that produced them.