Get Started
Research questionHow can policy optimization for long-horizon LLM agents preserve useful transitions across updates when rollout groups are small?Each policy update may discover useful transitions that later updates cannot use, while small rollout groups make relative advantage estimates noisy. This makes it difficult to attribute delayed outcomes to the actions and states that produced them.
AI Agents
LLM Pretraining & Post-training
Machine Learning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.TIGPO: Temporal Instance-Graph Policy Optimization for Long-Horizon LLM AgentsApplies to LLM-agent policy optimization. The source evaluates this setting on ALFWorld and WebShop and describes using persistent task-level transition histories and revisits as detached structural or statistical references, not as replayed policy-loss data.research paper · Sep 3, 2026
Related questions
How can long-horizon LLM agents learn when to group actions without overcommitting?How can LLM agents jointly adapt reasoning policies and hierarchical skill libraries during reinforcement learning?How can LLM agents generalize to unseen tasks without directly fine-tuning their policies?How can long-horizon LLM agents preserve answer quality under tight prompt-token budgets?
Home
Topics
Search
Library