Get Started
Home
Topics
Search
Library
Research questionHow can low-resource Thai TTS learn a fixed voice from synthetic speech while preserving pronunciation and prosody?With little speaker-specific data, a fixed-voice Thai synthesizer can avoid the inference cost of voice cloning, but synthetic targets may reproduce teacher errors or omit difficult text. Thai word boundaries, lexical tones, names, numbers, and code-switching make pronunciation and prosody particularly sensitive to these choices.
Audio & Speech
Audio & Speech Processing
Machine Learning
Natural Language Processing
Small / On-device Models
Latest papersRecent research connected to this question, newest first.Building and Evaluating Fixed-Voice Thai TTS from Synthetic SpeechThe study uses a large voice-cloning teacher conditioned on a short voice reference, such as 15 seconds, to generate all training speech for an 82M-parameter student that synthesizes without reference audio. Evidence is limited to the reported Thai and English error rates, keyword and pause accuracy, speaker similarity, speaking rate, and comparisons among the evaluated systems.research paper · Sep 3, 2026
Related questions
How can speech synthesis produce natural dubbing and full-duplex dialogue without forced alignment or explicit duration prediction?How can speech deepfake detectors focus on synthesis artifacts while generalizing to unseen speakers?How can speaker-attributed ASR identify who said what as speech arrives with low latency?How can spoken language detection adapt to underrepresented accents under low-resource constraints without overfitting?