Get Started
Research questionHow can offline goal-conditioned reinforcement learning learn reliable values for long-horizon tasks without compounding overestimation?In offline goal-conditioned reinforcement learning, long-range value estimates depend on shorter-range estimates that may already be inaccurate. Repeated max-based backups can amplify those errors across long-horizon tasks.
Machine Learning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Recursive Value Learning for Long-Horizon Offline Goal-Conditioned RLThe problem concerns offline goal-conditioned reinforcement learning for long-horizon goal-reaching tasks using observed trajectories. Evidence is reported across diverse tasks and five challenging OGBench tasks, so it does not establish performance beyond those settings.research paper · Sep 2, 2026
Related questions
How can offline reinforcement learning improve policies beyond dataset support while keeping value estimates reliable under distribution shift?How can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?How can reinforcement learning reliably satisfy Value-at-Risk constraints during policy training?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?
Home
Topics
Search
Library