Get Started
Home
Topics
Search
Library
Research questionHow can we evaluate voice-agent robustness when spoken task-oriented interactions are scarce and behaviorally diverse?Collecting enough spoken task-oriented interactions is expensive, while existing datasets often cover too few domains, speakers, or conversational behaviors. These limitations make it difficult to test whether voice agents handle the variability of real users.
AI
AI Agents
Audio & Speech
Audio & Speech Processing
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.SpokenUS: A Spoken User Simulator for Task-Oriented DialogueThe source introduces SpokenTOD, containing 52,390 dialogues and 1,034 hours of speech across diverse speakers and domains, augmented with cross-turn slots, barge-in, disfluency, and emotional prosody. It also presents SpokenUS, a spoken task-oriented user simulator with a dedicated turn-taking head; reported evidence includes goal coverage, human mean opinion scores, and analyses of the challenges its behaviors pose for voice agents.research paper · Sep 1, 2026
Related questions
How can full-duplex voice agents infer role-implied behavior while managing overlapping speech and conflicting instructions in real time?How can in-car LLM agents respond consistently to incomplete requests they cannot safely fulfill?How should conversational foundation models be evaluated when latency and generation speed shape user experience?How can safety evaluations measure harmful actions by computer-using agents rather than chatbot refusals?