Get Started
Home
Topics
Search
Library
Research questionHow can alignment systems infer the multiple criteria behind human pairwise preferences?Pairwise preference labels record which option people choose, but not the interacting considerations behind that choice. This makes it difficult to build preference models that are both faithful to judgments and interpretable.
AI
Alignment & Safety
Evaluation & Benchmarks
LLM Pretraining & Post-training
Reasoning
Latest papersRecent research connected to this question, newest first.Democratic ICAI: Debating Our Way to Steering Principles from PreferencesThe evidence concerns creative preference tasks evaluated on MuCE-Pref and LiTBench, using LLM-based and decision-tree judges and downstream constitution-induced preference labels. It compares structured rationale-based preference modeling with deliberative prompting and principle-based baselines, reporting preference-prediction and constitution-quality outcomes.research paper · Sep 3, 2026
Related questions
How can text-to-image diffusion models use multiple candidates and continuous rewards beyond pairwise preferences?How can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?How can subjective-label models remain calibrated across legitimate differences in human values?How can we measure an AI agent’s tacit alignment with a human when objectives, communication, and feedback are limited?