Research questionHow can video-language models capture the distribution of human interpretations of dynamic facial expressions?Facial-expression datasets often reduce genuine perceptual disagreement to one ground-truth label. This makes it difficult to determine whether a model can represent the range of interpretations that dynamic expressions elicit.