Get Started
Research questionHow can zero-shot voice conversion transfer an unseen speaker’s identity while preserving content in low-latency streaming?The converter must preserve linguistic content while reproducing the target speaker’s identity under tight latency constraints. Streaming and chunked processing can make it difficult to maintain natural, intelligible speech without degrading speaker similarity.
Audio & Speech
Audio & Speech Processing
Machine Learning
Latest papersRecent research connected to this question, newest first.X-VC: Zero-shot Streaming Voice Conversion in Codec SpaceThe evidence concerns X-VC, a one-step converter operating in pretrained neural-codec latent space with target reference speech. Experiments on Seed-TTS-Eval cover English and Chinese, including same-language and cross-lingual conversion, and report word error rate, speaker similarity, and real-time performance; the supplied evidence does not establish behavior beyond these settings.research paper · Sep 4, 2026
Related questions
How can streaming neural audio codecs preserve speech intelligibility under zero-lookahead, low-latency constraints?How can speaker-attributed ASR identify who said what as speech arrives with low latency?How can creators generate reusable multi-speaker voices and expressive audio scenes from instructions or reference recordings?How can zero-shot VLA transfer across unseen embodiments be assessed without confounding task or protocol differences?
Home
Topics
Search
Library