Research questionHow can conversational reinforcement learning coordinate strategic utterance choices with token generation under sparse, delayed rewards?A conversational agent must decide both which communicative strategy to pursue and how to realize it token by token. Feedback often arrives at the utterance or conversation level, making credit assignment across long interactions difficult.