Get Started
Home
Topics
Search
Library
Research questionHow can we select nonredundant preference data across safety datasets without losing robustness?Safety-alignment preference datasets can contain substantial redundancy, making it difficult to reduce training data without discarding directions that capture distinct safety risks. Selecting examples independently may also overlook relationships between shared safety patterns and dataset-specific residual risks.
Alignment & Safety
Evaluation & Benchmarks
LLM Pretraining & Post-training
Machine Learning
Latest papersRecent research connected to this question, newest first.DOG-DPO:Dynamic Optimization in Geometry for Safety AlignmentThe source concerns preference-pair selection before DPO training for large language models. Evidence covers six safety benchmarks and two model backbones, including selection at 11% of the preference data and reported utility–robustness results; it does not establish performance beyond those settings.research paper · Sep 2, 2026
Related questions
How can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?How can safety-critical control prioritize multiple uncertain risks while certifying how far a policy is from optimal?How can MORL policy selection account for behavioral differences hidden by similar objective values?How can we select visual instruction-tuning examples under a fixed budget while preserving alignment and dataset coverage?