Research questionHow can we evaluate voice-agent robustness when spoken task-oriented interactions are scarce and behaviorally diverse?Collecting enough spoken task-oriented interactions is expensive, while existing datasets often cover too few domains, speakers, or conversational behaviors. These limitations make it difficult to test whether voice agents handle the variability of real users.