Get Started
Home
Topics
Search
Library
Research questionHow can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?Offline preference optimization backpropagates through every chosen–rejected pair even though their usefulness changes as the policy evolves. Some pairs provide little information, while others can produce noisy or destabilizing updates.
AI
Alignment & Safety
LLM Pretraining & Post-training
Machine Learning
Reasoning
Latest papersRecent research connected to this question, newest first.Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning AlignmentThe source concerns reasoning models trained from fixed chosen–rejected pairs. Evidence comes from mathematical reasoning benchmarks across multiple model scales, with comparisons against full-data and size-matched baselines.research paper · Sep 3, 2026
Related questions
How can reinforcement learning post-training prioritize useful reasoning prompts as learning signals shift?How can offline reinforcement learning improve policies beyond dataset support while keeping value estimates reliable under distribution shift?How can alignment systems infer the multiple criteria behind human pairwise preferences?How can idle inference resources reduce scarce-GPU training cost without biasing gradient estimates?