Get Started
Home
Topics
Search
Library
Research questionHow can speech and facial cues be combined for emotion recognition when timing and class balance vary?Speech and facial expressions can provide complementary emotional evidence, but their signals may occur at different times and emotion classes may be unevenly represented. A recognition system must combine these sources without allowing temporal mismatch or dominant classes to distort predictions.
AI
Audio & Speech
Audio & Speech Processing
Computer Vision
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.Enhancing Multimodal Emotion Recognition via Multi-Feature Encoding and Attention-Based FusionThe source evaluates audio-visual emotion recognition using multiple speech representations and facial-sequence features, with learned feature fusion on the MELD and IEMOCAP datasets. Its evidence is limited to the evaluated tasks, datasets, models, and unbalanced-data settings.research paper · Sep 4, 2026
Related questions
How can conversational speech emotion recognition use cross-speaker context without confusing each speaker’s emotional trajectory?How can video facial-expression recognition personalize vision-language models under shifts without costly test-time optimization?How can ViT-based video facial-expression recognition detect subtle, localized temporal changes that global attention overlooks?How can real-time emotion recognition in VR handle upper-face cues occluded by head-mounted displays?