VoxPolyMem builds long-term agent memory for multi-party spoken conversations by tracking recurring speakers acoustically, storing who said what to whom as a directed graph, and training a retrieval agent with Evidence-Gain GRPO that rewards each round only for newly-retained supporting evidence.
Most LLM memory systems assume a two-party text or image-text chat: one user, one assistant, remember facts about the user. That assumption breaks for in-car assistants, meeting bots, and household devices, where several people speak, their roles matter, and the same person reappears across sessions with no login.
Two concrete failures follow. First, if you only store ASR transcripts, you lose the acoustic fingerprint that lets you recognize this is the same person who spoke last Tuesday. Second, prior memory systems like Mem0 and temporal-graph approaches like Zep store facts and timelines but not interaction roles: they can tell you “the meeting was rescheduled” but not “Alice told Bob it was rescheduled, while Carol was not addressed.” Personalized answering (“what did I agree to?”) and attribution (“who promised that?”) both need those roles.
The paper’s baselines are the standard memory stack (Mem0, A-MEM, LightMem) plus multimodal retrieval systems. None jointly handle recurring speaker ID, participant-scoped memory, and interaction attribution over spoken multi-session histories.
Three pieces fit together.
1. Online speaker identification. Each audio turn is embedded with a pretrained ECAPA-TDNN voice encoder. The system keeps a growing table of anonymous speaker_id → voiceprint entries. A new turn is matched by cosine similarity to the closest stored voice. Above a match threshold it joins that identity; above a stricter update threshold the stored voiceprint is nudged toward the new sample via exponential moving average. Below both, a new ID is created. After each session, fragmented IDs that likely belong to one person are conservatively merged. No participant roster or speaker count is required up front.
2. Interaction-aware hierarchical memory. Three layers sit on top of raw turns:
•
Interaction memory: a directed graph where edges record speaker --speaks--> message and message --addresses--> participant. This is the structural novelty, explicit addressee links, not just speaker tags.
•
Fact memory: self-contained statements extracted by an LLM from a short rolling window, each tagged with speaker, addressees, and time.
•
Participant profiles: name, aliases, associated voiceprint embeddings, background, and preferences.
Facts and profile fields keep refer_id pointers back to the source turns, which matters for scoring retrieval later.
3. Agentic retrieval trained with EG-GRPO. Answering is cast as a short sequential-decision loop (budget: 3 rounds). At each step the policy picks: which memory layer to hit, which tool (BM25, dense text, text-to-image, image-to-image), a rewritten query, and optional speaker/addressee filters. Results from multiple tools are fused with Reciprocal Rank Fusion, deduplicated, and the top ~15 kept as the retained context. An LLM then either answers or flags what is still missing.
Training uses Evidence-Gain GRPO, a variant of Group Relative Policy Optimization (GRPO). The key move: instead of rewarding only the final answer, each round scores a candidate action by how much new supporting evidence it pulls into the retained context, with a rank-discounted bonus for putting that evidence near the top. Evidence already retained on earlier rounds of the same branch gets zero credit, which pushes the policy toward complementary searches rather than repeated near-duplicates.
for t in range(budget): # budget = 3
action = policy(query, asker, retained_ctx, history)
hits = run_tool(action.layer, action.tool, action.rewrite, action.filters)
retained_ctx = select(rrf_dedup(retained_ctx + hits), k=15)
gained_ids = support_ids(retained_ctx) & gold_ids - seen_ids
reward = lam * len(gained_ids)/len(gold_ids) + (1-lam) * rank_discount(gained_ids)
seen_ids |= support_ids(retained_ctx)
if llm_says_sufficient(retained_ctx): break
answer = fixed_answer_model(query, retained_ctx)
Only the retrieval policy’s tokens (layer, tool, rewrite, filters) get gradient updates; the answer model stays frozen.
The authors also release VoxPolyBench: 18 scenarios, 176 sessions, 18.9 hours of synthesized speech, 1,527 QA pairs across memory evolution, personalization, retrieval/reasoning, and interaction attribution. Baselines get oracle speaker labels; VoxPolyMem does speaker ID itself.
•
Headline: VoxPolyMem scores 85.0 on VoxPolyBench vs 61.4 for the strongest external baseline (a 23.6-point gap). Personalization jumps from the 40s (baselines) to 76.1; interaction attribution from the 30s–60s to 87.5. These are the two subsets where explicit speaker/addressee edges should matter most, and they do.
•
Public benchmarks (image-text, not speech): VoxPolyMem reaches 89.6 on Mem-Gallery and 74.4 on H2HMem-Multi, beating the strongest baselines by 11.8 and 8.4 points. Suggests the hierarchy and agentic retrieval help beyond the authors’ own benchmark.
•
Ablations (built on the no-RL variant): removing fact memory and profiles drops H2HMem-Multi from 72.1 to 63.6; removing interaction edges drops VoxPolyBench from 84.0 to 78.6; forcing single-round retrieval is the worst single ablation. Removing query rewriting alone barely hurts, so iterative retrieval matters more than rewriting without RL.
•
Retrieval-policy comparison: against Search-R1 QA setting (terminal reward), terminal-coverage GRPO, and MoT-GRPO, EG-GRPO leads on answer score and recall on every benchmark and uses fewer retrieval rounds (1.2–1.3 vs 1.4–1.8). Notably, Search-R1 QA setting’s terminal reward actually degraded answer scores vs no RL in this setup.
•
Speaker tracker sanity check on real recordings: 95.0% attribution accuracy on IEMOCAP, 87.9% on AMI Meeting Corpus, given oracle utterance boundaries.
One honest wrinkle the authors flag: the LLM judges (GPT-4.1-mini + GPT-4.1) agree with human scores within ~1.8 points on average, but that’s aggregate agreement, not per-item.
•
If you’re building a shared-device assistant (car, home hub, meeting bot): the explicit addressee graph is the piece worth copying first. The ablations show it is what drives attribution and personalization gains, more so than the RL training. You can get most of the benefit without any policy learning, since the no-RL variant already beats every external baseline on average.
•
If you already have an agentic retrieval loop and are considering RL for it: the comparison against Search-R1 QA setting is the useful data point. Terminal answer rewards can underperform no-RL in a multi-tool, multi-layer setup. Per-round credit for new evidence (what EG-GRPO does) is worth testing before investing in trajectory-level reward shaping.
•
For evaluation: VoxPolyBench is worth knowing about if you work on voice assistants with persistent users, because existing speech benchmarks test within-session understanding or long-context recall but not cross-session identity plus participant-dependent memory. Code and data: voxpolymem.github.io.
•
Caveat on transfer: the speaker tracker was tested on IEMOCAP and AMI Meeting Corpus with reference utterance boundaries supplied. If your pipeline has to do its own voice-activity detection and diarization, those accuracy numbers are an upper bound.
•
VoxPolyBench is synthesized speech with planned events and LLM-generated dialogues, not real multi-party recordings. The authors are explicit that this does not establish robustness to real accents, overlap, noise, or genuinely spontaneous conversation.
•
Baselines on VoxPolyBench received ground-truth speaker labels while VoxPolyMem had to earn them, which cuts both ways: VoxPolyMem’s advantage is real but it is not a like-for-like comparison on retrieval alone.
•
Headline gains include the memory structure and the retrieval loop and RL. RL contributes about 1–3 points over the no-RL variant on each benchmark. Most of the lift is architectural, not from policy optimization.
•
Scoring uses LLM-as-judge (GPT-4.1 family) with a small human audit. The audit shows close mean agreement but does not establish item-level reliability.
•
Voiceprint storage plus persistent interaction graphs is a privacy surface the paper raises in its ethics statement. Deployment needs consent, retention limits, and deletion paths that the research artifact does not provide.