Get Started
Home
Topics
Search
Library
Research questionHow can assistants remember and reason about how users sounded across long, multi-session conversations?Transcripts preserve words but can discard emotion labels, prosody descriptors, and voice events. Assistants working across long, multi-session histories may therefore fail on questions whose answers depend on how a user spoke.
AI
AI Agents
AI Memory
Audio & Speech
Audio & Speech Processing
Evaluation & Benchmarks
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.VoiceLongMemEval: Do Assistants Remember How You Sounded?The evidence concerns VoiceLongMemEval, a benchmark for questions requiring paralinguistic information attached to conversational turns. It compares transcript-only models, models given text-based paralinguistic metadata, and audio-native models that extract cues directly from speech; the reported results also indicate that standard ASR pipelines discard this signal. The evidence is benchmark-based rather than a deployment evaluation.research paper · Sep 2, 2026
Related questions
How can long-term conversational QA agents retrieve and reason over temporally dispersed dialogue history?How can conversational agents retrieve the right memories when users rely on implicit conversational context?How can multimodal agents maintain consistent person identities and reason about relationships across long video memories?How can personalized LLM agents retrieve time-valid memories of persistent and evolving user states?