Research questionCan multimodal models match human judgments of facial attractiveness, not merely rank faces correctly?A model can track which faces people prefer while still assigning scores that are systematically too high and too compressed. Agreement in rankings therefore does not establish that its attractiveness ratings reflect human judgments in absolute terms. Latest papersRecent research connected to this question, newest first.Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractivenessEvidence comes from a preregistered exploratory comparison of 2,513 human participants with Claude, Gemini, GPT, and Grok. The models generally overrated faces, used a narrower rating range, correlated strongly with human rankings, and differed in cues associated with age, ethnicity, and gender; Grok showed the lowest agreement with humans.research paper · Sep 2, 2026