Research questionHow can subjective-label models remain calibrated across legitimate differences in human values?Subjective annotations can reflect legitimate differences in human values rather than mere labeling noise. A model may therefore appear accurate overall while remaining miscalibrated for particular value groups.