Get Started
Home
Topics
Search
Library
Research questionHow can full-duplex dialogue models learn natural acoustic turn-taking without degrading semantic responses?Full-duplex systems must decide when to listen, speak, or yield while generating semantically appropriate responses. Synthetic text supervision misses the fine-grained acoustic timing of human conversation, while changing turn-taking behavior can disrupt semantic capability.
AI
Audio & Speech
Audio & Speech Processing
LLM Pretraining & Post-training
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex DialogueThe source studies a neural finite-state-machine framework that serializes turn-taking control and response generation on one causal tape. Its reported approach transforms real human–human spoken dialogues into turn-taking supervision and uses configurable human–agent text dialogues for semantic behavior; experiments report improved turn-taking and recovered foundation-model semantic capability.research paper · Sep 3, 2026
Related questions
How can speech language models consistently use paralinguistic cues in open-ended, multi-turn dialogue?How can speech synthesis produce natural dubbing and full-duplex dialogue without forced alignment or explicit duration prediction?How can full-duplex voice agents infer role-implied behavior while managing overlapping speech and conflicting instructions in real time?How can conversational reinforcement learning coordinate strategic utterance choices with token generation under sparse, delayed rewards?