Get Started
Home
Topics
Search
Library
Research questionHow can deployed vision models resist diverse backdoor triggers without model internals or auxiliary data?Backdoor triggers can cause targeted misclassification while leaving clean accuracy high, making compromised image regions difficult to identify from ordinary predictions. Defenders without model internals, training data, or clean validation images must distinguish malicious content from benign visual information.
AI
Alignment & Safety
Computer Vision
Image & Video Processing
Machine Learning
Latest papersRecent research connected to this question, newest first.Beyond Small Patches: Black-Box Detection and Purification of Diverse Backdoor TriggersApplies to trained vision models accessed as black boxes, without model internals, training data, or clean samples. The source reports experiments on blended, sparse, varying-size, and multiple triggers across diverse datasets; its evidence concerns inference-time mitigation and does not establish behavior beyond those experiments.research paper · Sep 2, 2026
Related questions
How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can black-box adversarial attacks disrupt image-based continuous-control agents with fewer environment queries?How can one vision model remain robust across changing and unseen adversarial perturbation budgets?How can language models distinguish trusted instructions from untrusted text to resist prompt injection?