Get Started
Research questionHow can reinforcement learning reliably satisfy Value-at-Risk constraints during policy training?Tail-risk constraints such as Value-at-Risk can be difficult to estimate and enforce stably during policy updates, particularly with tight violation thresholds and dense costs.
Machine Learning
Reinforcement Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Cantelli Constrained Policy OptimizationApplies to constrained reinforcement learning with Value-at-Risk constraints. The source provides a moment-based Cantelli bound, trust-region training guarantees, and empirical results across tested environments; evidence is limited to those theoretical and experimental settings.research paper · Sep 10, 2026
Related questions
How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?How can online reinforcement learning make VLA manipulation precise without value drift or prohibitive cost?How can group-relative policy optimization enforce constraints without normalization coupling reward and constraint objectives?How can offline goal-conditioned reinforcement learning learn reliable values for long-horizon tasks without compounding overestimation?
Home
Topics
Search
Library