Get Started
Home
Topics
Search
Library
Research questionHow can online reinforcement learning train multi-turn computer-use agents under partial observability, sparse rewards, and costly rollouts?In a long desktop task, each action changes the next observation and available actions, while success may be revealed only at termination. Slow environment feedback makes collecting and coordinating enough interactive trajectories difficult.
AI
AI Agents
Evaluation & Benchmarks
Machine Learning
Multimodal Models
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use AgentsEvidence comes from executable sandbox environments and includes OSWorld-Verified performance. It supports claims about online training in that benchmarked setting, not unrestricted real-world desktop use or tasks without verifiable outcomes.research paper · Sep 4, 2026
Related questions
How can terminal-agent training environments stay challenging as models improve without costly on-policy synthesis?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?How can reinforcement-learning agents learn robust ad-hoc teamwork without pre-trained partners or hand-tuned partner generation?How can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?