Research questionHow can reinforcement-learning data acquisition be optimized when interactions are costly, slow, and human-mediated?In business and healthcare operations, reinforcement-learning interactions can be slow, expensive, and require human involvement. Limited interaction data makes it difficult to select effective policies efficiently and quantify the value of additional acquisition.