Research questionHow can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?Offline preference optimization backpropagates through every chosen–rejected pair even though their usefulness changes as the policy evolves. Some pairs provide little information, while others can produce noisy or destabilizing updates.