Research questionHow can speech synthesis produce natural dubbing and full-duplex dialogue without forced alignment or explicit duration prediction?Dubbing and full-duplex conversation require timing to emerge alongside voice, prosody, and interaction rather than from precomputed alignments. Without explicit durations, the generator must still coordinate text with natural turn-taking and sustained audio output.