Research questionHow should subtoken granularity be chosen to reduce masked diffusion language-model training loss?Masked diffusion language models operating on subtokens can incur higher cross-entropy loss with BPE-based tokenizers. The choice of subtoken granularity also lacks clear guidance, making it difficult to relate tokenizer structure to training and downstream behavior.