VoxMem is a benchmark for spoken-conversation memory that separates what must be remembered (words, speaker, vocal style, background sound) from how it’s used (retrieval, cross-session reasoning, temporal tracking, refusal). Across 15 audio LLMs, none exceeds 40% at 32K-token histories.
If you build a voice assistant that holds many conversations with a user over weeks, it needs to recall more than transcripts. Who spoke, how they sounded, and what was audible in the background can all matter for the next answer. A user might ask “what did I say last time I was in the car?” where “in the car” is only recoverable from engine noise, not words.
Existing spoken-memory benchmarks mostly test lexical recall from a single long recording or a single dialogue. They don’t jointly probe non-lexical acoustic evidence and multi-session structure. So when a model fails, you can’t tell whether it lost the word, the voice, the tone, or the ability to combine evidence across sessions. The authors argue you need a taxonomy that crosses evidence-type with memory-operation, and histories that span distinct sessions with gaps.
The closest prior efforts (Audio MultiChallenge, Vox-Infinity, and others) cover only a few cells of this grid, use ad-hoc operations, and don’t hold the question fixed as history grows, so length effects are confounded with question difficulty.
VoxMem is a dataset and evaluation protocol, not a model. The contribution is a two-axis taxonomy and a construction pipeline that keeps the question fixed while varying history length.
•
Evidence axis (4 types): speech semantics (words), speaker identity (who), paralinguistic cues (how, e.g., hesitation, laughter), environmental sound (what was audible).
•
Operation axis (4 types): Information Extraction (IE), Multi-Session Reasoning (MSR), Temporal Evolution Tracking (TET), and Answer Refusal (AR). AR items are derived by stripping the critical evidence from an answerable item, so the model should say “I don’t know.”
Each item is built in three stages. First, a structured plan fixes the question, gold answer, and which sessions carry evidence. Second, dialogues are written with user and assistant generated independently, so the assistant text can’t leak the acoustic answer. Third, user turns are synthesized with Higgs-TTS-3 using fixed VCTK voices per persona; paralinguistic cues come from TTS style controls, and environmental sounds from ESC-50 mixed at 10 dB SNR.
Each item is then embedded in four nested histories at 8K, 16K, 32K, and 64K tokens (measured on a shared Whisper encoder scale, roughly 2.5 to 20 minutes of audio). Longer histories strictly extend shorter ones by adding haystack sessions (same topic, wrong answer) and filler sessions (unrelated chat drawn from InstructS2S). The question and evidence are byte-identical across lengths, so any accuracy drop is attributable to surrounding context, not question difficulty.
for item in questions:
plan = make_plan(operation, evidence_type, answer)
evidence = write_dialogues(plan) # user/assistant separate
audio = synthesize(evidence, voices, cues, env_sounds)
quality_control(audio, plan) # drop or revise
for budget in [8K, 16K, 32K, 64K]:
haystack = pick_distractors(plan, budget)
filler = pad_from_instructs2s(budget)
history[budget] = interleave(evidence, haystack, filler)
ar_variant = strip_evidence(item) # paired refusal item
Quality control uses Gemini-3.7-Flash (which is not among the evaluated models) to verify answer uniqueness, that audio-native questions truly can’t be answered from a transcript, and that no haystack or filler accidentally supplies the answer.
The final benchmark has 3,196 instances over 34,743 sessions (177 hours), evaluated on 15 audio LLMs including open-weight (Qwen3-Omni, Baichuan-Audio, Audio-Flamingo-Next, Ultravox, Phi-4-Multimodal, Gemma, MiniCPM-o, FireRedAudio, MiMo-Audio) and proprietary systems (three Gemini variants, two Qwen variants).
•
No model exceeds 40% overall accuracy at 32K; the best reaches 38.5%. Proprietary models average 33.0%, open-weight 21.9%.
•
Words are remembered much better than everything else. At 32K, proprietary models average 55.6% on speech semantics versus 32.7% on speaker, 20.0% on paralinguistic, 21.9% on environmental. Open-weight shows the same ordering at lower absolute levels.
•
The acoustic-necessity filter works. On audio-native questions, giving the judge a perfect transcript instead of the audio drops accuracy from 46.2% to 4.4%, while speech-semantics items barely move (75.9% to 71.0%). So the audio-native items genuinely require audio.
•
Operation interacts with evidence type. Temporal Evolution Tracking (TET) on speech semantics reaches 44.5%, but collapses to 3.4% on paralinguistic and 1.2% on environmental. Models can track states expressed in words but not states carried by voice or ambient sound.
•
Refusal accuracy is inversely patterned. Models “abstain” more on paralinguistic (34.3%) and environmental (34.2%) questions than on semantics (19.9%). The authors read this as models bailing out when they can’t process the audio, not genuine awareness of missing evidence.
•
Longer context erodes access to the same evidence. From 8K to 64K, models retain 70.3% of their speech-semantics accuracy but only 63.2% of environmental. Because the question is held fixed, this isolates the length effect.
•
Error profiles differ by type. Speaker errors are mostly binding failures (wrong voice attached to right content, 48%). Paralinguistic errors are mostly localization failures (the cue never gets retrieved, 63%). So the three audio-native types fail for different underlying reasons.
•
If you’re evaluating a voice assistant for long-term memory, don’t rely on transcript-based benchmarks or single long recordings. Test speaker, paralinguistic, and environmental recall separately, because strong word-level scores mask large gaps elsewhere. VoxMem is the first released benchmark that does this across sessions with controlled length.
•
If you’re a model developer, the finding that TET collapses on acoustic states (3.4% paralinguistic, 1.2% environmental) suggests investing in stateful acoustic representations rather than only expanding context windows. The paper shows context expansion alone doesn’t fix access to information already in the window.
•
If you’re interpreting a model’s refusals as calibration, be careful. The paper’s refusal profile suggests high abstention on acoustic questions reflects inability to hear, not epistemic awareness. Worth testing on your own stack whether refusals correlate with actual missing evidence.
•
For benchmark designers, the construction trick of holding question and evidence fixed across nested context budgets is worth copying. It separates length sensitivity from question difficulty, which prior spoken-memory benchmarks confounded.
•
Assistant turns are provided as text, not speech. This was done so models that only consume audio (not generate it) can be evaluated identically, but it means the benchmark doesn’t test memory over assistant speech.
•
Only 10 of 15 models completed the 64K setting, so 64K comparisons use a reduced model set.
•
The LLM judge is Gemini-3.7-Flash, and three evaluated models are also Gemini. Cross-validation with GPT-5.6-Luna shows 0.88 pp favoritism toward Gemini outputs, small but not smaller than every pairwise gap at 8K.
•
Paralinguistic cues come from TTS style controls, not natural speech, so results reflect synthesized rather than spontaneous prosody. Environmental sounds are mixed at a single 10 dB SNR.
•
Some model version strings in the paper’s model table (e.g., “Gemini-3.8-Flash”, “GPT-5.6-Luna”, “Gemma-4”) appear to be forward-dated or non-standard; the paper lists them as-is and the specific checkpoints are recorded in the released code.