Research questionHow can zero-shot voice conversion transfer an unseen speaker’s identity while preserving content in low-latency streaming?The converter must preserve linguistic content while reproducing the target speaker’s identity under tight latency constraints. Streaming and chunked processing can make it difficult to maintain natural, intelligible speech without degrading speaker similarity.