Get Started
Home
Topics
Search
Library
Research questionHow can preference optimization focus supervision on differing entity slots when contrastive candidates share templates?Template-based contrastive candidates may share nearly all of their wording while differing in only a few entity slots. Sequence-level objectives can spread preference supervision across this shared content instead of emphasizing the distinctions that determine which candidate is preferred.
AI
Alignment & Safety
LLM Pretraining & Post-training
Machine Learning
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.KARMA: Knowledge graph-based Automated Reasoning Materialization and AlignmentThe research uses schema-constrained paths from biomedical, computer-science, and chemistry knowledge graphs to generate slot-aligned verbalizations for contrastive preference training. Its evidence covers those benchmark domains and comparisons with base language models, same-data supervised fine-tuning, sequence-level preference methods, and token-level preference methods; it does not specify deployment or access constraints.research paper · Sep 3, 2026
Related questions
How can preference optimization align whole-image preferences with token-specific content across spatial and denoising coordinates?How can offline preference optimization identify which chosen–rejected pairs merit gradients without destabilizing reasoning-model training?How can text-to-image diffusion models use multiple candidates and continuous rewards beyond pairwise preferences?How can language-model attention remain reliable beyond its training context?