Get Started
Research questionHow can reinforcement learning discover diverse high-reward states when long horizons hinder credit assignment and exploration?Long trajectories make it difficult to identify which actions led to reward and to explore enough of the action space to find diverse high-reward states. These difficulties are especially acute when many distinct objects or modes must be discovered.
Machine Learning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Action abstractions for amortized samplingThe source addresses entropy-seeking RL, particularly GFlowNets, whose policies sample from structured distributions. It studies action abstractions formed by chunking recurring subsequences from high-reward trajectories, with empirical evidence from synthetic and real-world environments showing improved sample efficiency on harder exploration problems and interpretable abstractions.research paper · Sep 2, 2026
Related questions
How can hierarchical reinforcement learning use incrementally acquired knowledge for long-horizon exploration with sparse rewards?How can reinforcement-learning data acquisition be optimized when interactions are costly, slow, and human-mediated?How can world-model reinforcement learning produce reliable long-horizon driving policies amid interactive traffic and diverse driving styles?How can reinforcement learning post-training for diffusion models avoid objective mismatch and high-variance updates?
Home
Topics
Search
Library