Research questionHow can reinforcement learning discover diverse high-reward states when long horizons hinder credit assignment and exploration?Long trajectories make it difficult to identify which actions led to reward and to explore enough of the action space to find diverse high-reward states. These difficulties are especially acute when many distinct objects or modes must be discovered.