Get Started
Home
Topics
Search
Library
Research questionHow can text-to-image diffusion models erase unwanted concepts while preserving benign concepts and resisting re-emergence?Editing out a target concept can damage related benign concepts, while prompts that exploit the model's denoising behavior may cause the target to reappear. Effective erasure therefore requires defining the target broadly enough for robustness without unnecessarily altering neighboring concepts.
AI
Alignment & Safety
Diffusion Models
Evaluation & Benchmarks
Image Generation
Latest papersRecent research connected to this question, newest first.To Erase, or Not to Erase: Robust Training-Free Concept Erasure with Preservation aware Adaptive Ranked Subspace ExpansionThe source presents a training-free editing approach for latent diffusion models that adapts erased and retained concept directions and searches for re-emergence triggers. It evaluates NSFW content, artistic styles, and objects under experimental attacks and utility tests; the evidence is limited to editable latent models and the tested concepts and attacks, rather than models available only through ordinary generation APIs.research paper · Sep 4, 2026
Related questions
How can text-guided image editing explain edits while keeping provenance resistant to white-box removal?How can text-to-image models preserve variation in unspecified visual factors under long, semantically dense prompts?How can text-to-image diffusion models preserve prompt alignment without the extra sampling cost of classifier-free guidance?How can audio-video diffusion models preserve intended conditioning when biased cross-modal attention reroutes semantics?