MemAdapter is a post-retrieval wrapper that decides how much each retrieved memory should influence an agent’s answer, by first imagining alternative tasks where that memory would mislead, then calibrating its role against the current query and evidence.
Say you’ve built a personal assistant with long-term memory. It remembers your friend is a cardiologist, that they prefer treatment A, that they dislike agile methodologies. Later, when the agent helps pick a treatment for a specific patient, it recommends treatment A. Not because the clinical evidence supports it, but because that’s what the user prefers. The memory is accurate. It’s relevant. And it still corrupted the answer.
This is memory-induced Sycophancy, and the paper’s preliminary study makes the problem concrete: on a controlled set of 75 cases where retrieved memories were manually verified as accurate and task-relevant, adding memory dropped accuracy from 96.0% to 77.3%. Of answers that were correct without memory, 19.4% flipped to wrong once memory was injected, and every flip aligned with the retrieved memory.
Prior fixes almost all intervene before reasoning: filter noisy memories at extraction (Recalling Too Well), gate them at retrieval (MemGate), or restructure them by domain. These assume the problem lives in the memory content. The paper’s point is that it doesn’t. A correct memory can still corrupt reasoning when invoked in the wrong context.
MemAdapter sits after retrieval and asks a different question: for each memory the retriever handed us, what is it actually allowed to justify in this task? Three stages, all LLM-prompted (no training).
Stage 1: Counterfactual Induction. Take each retrieved memory in isolation, without seeing the current query. Imagine several plausible tasks where this memory might show up. For each imagined task, write down what the memory can legitimately support, what it can only influence partially, and what it absolutely cannot justify. Compare across these hypothetical contexts to extract a boundary card: a condition-to-influence rulebook for this memory.
Stage 2: Context-Aware Reflection. Now bring in the real query and evidence. For each memory, consult its boundary card and decide the task-specific role: what it may support, which parts of the response it may shape, how strongly, and where its influence must stop. Output a natural-language use instruction per memory.
Stage 3: Evidence-Based Reasoning. Generate the final answer under those instructions. Internally, each major response span gets tagged with its source and the applicable instruction, and the model verifies that memory use stayed within scope before returning.
boundaries = []
for m in retrieved_memories:
hypothetical_tasks = propose_task_types(m)
assessments = [assess_use(m, h) for h in hypothetical_tasks]
boundaries.append(consolidate_rules(assessments))
instructions = reflect(query, evidence, retrieved_memories, boundaries)
answer, trace = generate_with_grounding(query, evidence, retrieved_memories, instructions)
verify_each_span_against(trace, instructions)
return answer
The key design choice: Stage 1 is memory-conditional, task-blind. The boundary card is derived without seeing the query, so it’s a general-purpose risk profile. Stage 2 then resolves that prospective profile against the actual task.
Evaluated on three benchmarks that specifically probe memory misuse: MemSyco-Bench, PersistBench, and MemTrapBench. Tested across five memory backends (A-MEM, Mem0, NaiveRAG, MemoryBank, LightMem) with DeepSeek-V4-flash as the main backbone, and also with GPT-5.6-sol and qwen3-8b for cross-backbone checks.
•
Sycophancy reduction is the headline. On PersistBench’s sycophancy failure rate (3 attempts), MemAdapter cut failures from the 76-82% baseline range down to 51-58% across all five memory systems. The strongest competing intervention didn’t come close on this metric.
•
It evens out weak memory systems. On MemSyco-Bench average accuracy with DeepSeek, direct-generation scores ranged from 37.0 (LightMem) to 71.7 (A-MEM), a 34.7-point spread across backends. With MemAdapter, all five land in the 85.5-90.5 range, a spread of under 5 points.
•
Ablations isolate each stage’s contribution. Adding just Counterfactual Induction (stage 1) to NaiveRAG lifts average accuracy from 70.25 to 84.26. Adding Context-Aware Reflection brings it to 85.16. Adding Evidence-Based Reasoning takes it to 89.94. Each stage contributes, with the counterfactual induction stage carrying the largest single jump.
•
One real trade-off: on PersistBench’s beneficial memory failure rate (memories that should have been used), MemAdapter sometimes gets worse, going as high as 9% failure on MemoryBank and LightMem versus 3-5% baseline. Being more cautious about memory influence occasionally means ignoring memory that would have helped.
•
Latency is comparable to other post-retrieval interventions. Mean per-sample time is ~6.6-7.0 seconds versus ~6.0-7.1 for the baselines, with similar P95 and P99.
If you’re building a memory-augmented agent and have seen it defer to the user when it shouldn’t, the mechanism here is worth replicating even without the full framework. The core move is cheap: before generating, prompt the model to write a per-memory instruction specifying what each memory may and may not justify. The paper’s prompts are in Appendix H and the code is at GitHub.
If you’re deciding where to spend engineering effort on memory reliability, this paper argues for the post-retrieval stage. The preliminary study shows that filtering noisy memories at extraction won’t catch the failure mode, because accurate memories cause it too. Intervention after retrieval, operating on the retrieved set as given, is a complementary layer rather than a replacement for existing filters.
One prerequisite to watch: MemAdapter spends extra LLM calls on each memory (counterfactual induction runs per-memory). If your retrieval returns 10 memories per query, that’s at least 10 extra LLM round-trips before you even start answering. The paper doesn’t break out token cost, only wall-clock latency, but token spend scales linearly with retrieved-memory count.
Worth testing before deployment: whether the beneficial-memory regression matters for your use case. If your agent’s job is mostly to actively surface and apply user preferences (concierge-style personalization), the increased caution may hurt more than the sycophancy reduction helps. If it’s answering factual or decision-support questions where user beliefs leak in as pseudo-evidence, the trade is probably favorable.
•
Every component is LLM-prompted. No model is trained or fine-tuned. The entire framework’s quality rides on the backbone’s ability to do counterfactual reasoning and self-reflection reliably. Weaker open models may degrade the boundary cards into noise, though the Qwen3-8B results suggest it still helps at that scale.
•
The three benchmarks were all built specifically to study memory misuse, and two of them share authors with this paper. Independent evaluation on broader agent tasks would strengthen the generalization story.
•
The beneficial-memory trade-off is real and memory-system-dependent. The paper reports it transparently but doesn’t offer a mechanism to tune the caution/helpfulness dial.
•
Added latency and token cost scale with the number of retrieved memories. For large retrieval budgets, the per-memory counterfactual induction stage will dominate.