Get Started
Research questionCan black-box attackers identify and reconstruct prompts supposedly removed from language models without knowing them in advance?Machine unlearning may suppress responses to removed data without eliminating signals that reveal what was removed. The difficulty is determining whether an attacker can use those signals to discover and reconstruct forgotten prompts that are initially unknown.
AI
Alignment & Safety
LLM Pretraining & Post-training
Latest papersRecent research connected to this question, newest first.Extracting Forgotten Prompts from Targeted Unlearned ModelsEvidence comes from experiments involving NPO, DPO, and LUNAR across three datasets and three large language models, using a query-limited black-box attack. The reported recovery results apply to those methods, datasets, models, and access conditions, and do not establish behavior for other unlearning procedures or model-access regimes.research paper · Sep 3, 2026
Related questions
How can black-box systems detect and mitigate reward hacking in self-evolving language-model loops?How can class unlearning be audited for recoverable forget-class structure with only white-box model access?How can language models distinguish trusted instructions from untrusted text to resist prompt injection?How can retrieval-augmented language models resist ordinary-looking GEO-optimized documents that distort synthesized answers?
Home
Topics
Search
Library