Get Started
Home
Topics
Search
Library
Research questionHow can discrete audio tokenizers preserve semantics and acoustic fidelity for both understanding and generation?Audio tokenizers must compress continuous sound into manageable discrete sequences without losing either high-level content or acoustic detail. Separate semantic and acoustic streams can also introduce redundancy or misalignment when the same representation must support analysis and synthesis.
AI
Audio & Speech
Audio & Speech Processing
Research Paper
Sound
Technology
Latest papersRecent research connected to this question, newest first.EntangleCodec: A Unified Discrete Audio Tokenizer via Semantic-Acoustic EntanglementThe source concerns a unified discrete audio tokenizer for audio-language models, covering speech, music, and general audio. It reports evidence on reconstruction, audio understanding, text-to-speech, text-to-audio generation, and model scaling, using caption-aligned semantic-acoustic representations and a flow-matching diffusion decoder.research paper · Sep 3, 2026
Related questions
How can sparse autoencoders capture language-model features that persist across token sequences?How can 3D tokenizers preserve reconstruction fidelity with extremely short token sequences?How can speech enhancement use continuous audio representations while balancing quality against decoding cost?How can residual quantization encode semantic IDs while preserving graded similarity and directional residual variation?