Get Started
Research questionHow can conversational speech emotion recognition use cross-speaker context without confusing each speaker’s emotional trajectory?Emotion in dialogue depends on both a speaker’s own prior turns and what other participants say. Treating every adjacent utterance as one sequence can blur these distinct sources of temporal evidence across different time scales.
Audio & Speech
Audio & Speech Processing
Machine Learning
Latest papersRecent research connected to this question, newest first.Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in ConversationThe source describes a system using self-supervised speech representations, frame- and dialogue-scale state-space modeling, and independent speaker-wise dynamic CRF chains. Cross-speaker turns affect contextual emotion scores without becoming transitions in another speaker’s chain. Evidence is reported for audio-only recognition on IEMOCAP and MELD, with matched controls examining speaker-wise factorization and CRF modeling.research paper · Sep 4, 2026
Related questions
How can speech and facial cues be combined for emotion recognition when timing and class balance vary?How can assistants remember and reason about how users sounded across long, multi-session conversations?How can visual speech recognition resolve ambiguous words without committing before enough context is available?How can creators generate reusable multi-speaker voices and expressive audio scenes from instructions or reference recordings?
Home
Topics
Search
Library