Research questionHow can omni-modal models generate speech and temporally coordinated 3D facial animation?Semantic reasoning in language models produces relatively discrete representations, while facial animation requires dense, temporally precise motion coordinated with speech. Bridging these different granularities is difficult when the model must generate both modalities together.