Research questionHow can online reinforcement learning train multi-turn computer-use agents under partial observability, sparse rewards, and costly rollouts?In a long desktop task, each action changes the next observation and available actions, while success may be revealed only at termination. Slow environment feedback makes collecting and coordinating enough interactive trajectories difficult.