Get Started
Home
Topics
Search
Library
Research questionHow can video-language models capture the distribution of human interpretations of dynamic facial expressions?Facial-expression datasets often reduce genuine perceptual disagreement to one ground-truth label. This makes it difficult to determine whether a model can represent the range of interpretations that dynamic expressions elicit.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Multimodal Models
Latest papersRecent research connected to this question, newest first.Chehre: An Emoji-Prompted Dataset to Explore Perceptual Flexibility in Video Language ModelsThe source studies this task with Chehre, a privacy-preserving dataset of 2,111 synthetic-face videos based on 40 facial-expression prompts. Each video has roughly 30 annotations from a pool of 1,242 perceivers; the evidence evaluates selected video-language models and examines persona prompting as a way to shift model perception. The findings are limited to this dataset, task, and evaluated models.research paper · Sep 2, 2026
Related questions
How can video facial-expression recognition personalize vision-language models under shifts without costly test-time optimization?How can sign-language production models generate grammatical non-manual features without losing expressive variation?How can streaming video-language models cut frame-encoding latency while preserving question-relevant evidence?How can ViT-based video facial-expression recognition detect subtle, localized temporal changes that global attention overlooks?