Get Started
Research questionHow can reinforcement-learning data acquisition be optimized when interactions are costly, slow, and human-mediated?In business and healthcare operations, reinforcement-learning interactions can be slow, expensive, and require human involvement. Limited interaction data makes it difficult to select effective policies efficiently and quantify the value of additional acquisition.
Machine Learning
Reinforcement Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Optimal Data Acquisition for Reinforcement Learning: A Large Deviations PerspectiveThe source develops a theoretical efficiency measure based on the exponential decay of policy-selection error, characterizes optimal acquisition through large deviations, and studies a tractable relaxation with an adaptive acquisition algorithm. It also extends the framework to linear function approximation and reports supporting numerical experiments.research paper · Sep 4, 2026
Related questions
How can distributed VLA reinforcement learning coordinate variable-latency simulation, inference, and optimization?How can we learn near-optimal policies from transition samples in large or infinite-state-action MDPs using function approximation?How can reinforcement learning discover diverse high-reward states when long horizons hinder credit assignment and exploration?How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?
Home
Topics
Search
Library