Get Started
Home
Topics
Search
Library
Research questionHow can subjective-label models remain calibrated across legitimate differences in human values?Subjective annotations can reflect legitimate differences in human values rather than mere labeling noise. A model may therefore appear accurate overall while remaining miscalibrated for particular value groups.
AI
Alignment & Safety
Machine Learning
Natural Language Processing
Statistical Machine Learning
Latest papersRecent research connected to this question, newest first.Labels have Human Values: Value Calibration of Subjective TasksThe evidence concerns NLP models trained on subjective labels across toxic chatbot conversations, value reasoning, text-to-image safety, and preference-alignment tasks. It covers binary, ordinal, and preference learning, with value groups identified from rationale similarity, expert taxonomies, or annotator sociocultural descriptors.research paper · Sep 3, 2026
Related questions
How can machine learning learn from contested labels without assuming one objectively correct target?How can active preference learning obtain scalable, calibrated uncertainty for neural reward models without full Bayesian inference?How can alignment systems infer the multiple criteria behind human pairwise preferences?How can evaluators distinguish missing knowledge from miscalibrated outputs in language models?