Get Started
Home
Topics
Search
Library
Research questionHow can speech language models consistently use paralinguistic cues in open-ended, multi-turn dialogue?Speech conveys information through tone, speaker characteristics, and background conditions as well as words. Models may recognize these signals yet fail to let them shape responses reliably when dialogue continues or instructions compete.
AI
Audio & Speech
Audio & Speech Processing
LLM Pretraining & Post-training
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.ParaBridge: Bridging Paralinguistic Perception and Dialogue Behavior in Speech Language ModelsThe source studies scaffold-free spoken-dialogue behavior after training, reporting results on Qwen3-Omni-thinking and another speech-language-model backbone. Evidence includes safety- and empathy-oriented dialogue settings, unseen cues, multi-turn context, and general capability checks; it does not establish behavior for systems or access conditions not described here.research paper · Sep 3, 2026
Related questions
How can full-duplex dialogue models learn natural acoustic turn-taking without degrading semantic responses?How can multimodal models resist harmful intent unfolding across multi-image, multi-turn conversations?How can multimodal models rely on images or audio rather than language shortcuts?How reliably do vision-language models correct repeated visually grounded false premises across dialogue turns?