Get Started
Research questionHow can offline reinforcement learning improve policies beyond dataset support while keeping value estimates reliable under distribution shift?Offline RL learns from a fixed dataset, so policy changes toward unsupported actions can expose the critic to distribution shift and unreliable value estimates. Staying too close to observed behavior, however, can limit meaningful policy improvement.
Machine Learning
Reinforcement Learning
Latest papersRecent research connected to this question, newest first.Multi-step Proximal Policy Improvement in Offline Reinforcement LearningThe source addresses offline actor updates using critic-defined policy improvement, with deterministic and diagonal-Gaussian policy settings. Evidence includes D4RL benchmark experiments and diagnostics examining refinement behavior and limitations caused by critic error.research paper · Sep 3, 2026
Related questions
How can offline goal-conditioned reinforcement learning learn reliable values for long-horizon tasks without compounding overestimation?How can non-incremental tree learners handle distribution shift in online self-play reinforcement learning?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?How can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?
Home
Topics
Search
Library