Research questionHow can full-duplex dialogue models learn natural acoustic turn-taking without degrading semantic responses?Full-duplex systems must decide when to listen, speak, or yield while generating semantically appropriate responses. Synthetic text supervision misses the fine-grained acoustic timing of human conversation, while changing turn-taking behavior can disrupt semantic capability.