Get Started
Home
Topics
Search
Library
Research questionHow can active preference learning obtain scalable, calibrated uncertainty for neural reward models without full Bayesian inference?Active preference learning must choose which comparisons to request, but reliable uncertainty estimates become expensive for neural reward models when inference considers all parameters. Poorly calibrated uncertainty can lead to less informative queries and inefficient reward learning.
AI
Alignment & Safety
Machine Learning
Reinforcement Learning
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Subspace Inference Enables Efficient Active Reward Learning from PreferencesThe evidence concerns sequential preference learning for neural reward models, evaluated on the D4RL and V-D4RL benchmarks. Reported outcomes include sample efficiency, runtime, scalability, calibration, and downstream offline reinforcement-learning policy performance; behavior beyond these benchmarks is not established.research paper · Sep 3, 2026
Related questions
How can RLVR reduce the cost of on-policy rollouts and reliable targets without hurting reasoning quality?How can multiple-choice music audio-language models estimate uncertainty well enough to abstain without costly ensembles or retraining?How can leaderboards quantify uncertainty from sparse, unevenly sampled pairwise LLM judgments?How can reward shaping reduce reward hacking in RLHF when rewards imperfectly capture human preferences?