Get Started
Home
Topics
Search
Library
Research questionHow can vision-language models correct unsafe generations token by token without disrupting safe reasoning?Always-on safety mechanisms can alter decoding even when a generation is already safe, potentially degrading the model’s general multimodal capabilities. The challenge is to identify unsafe generation states as they arise and intervene without disturbing safe trajectories.
AI
Alignment & Safety
LLM Pretraining & Post-training
Multimodal Models
Latest papersRecent research connected to this question, newest first.SafeRI: Recognition and Intervention for Token-Level Safety Intervention in Large Vision Language ModelsThe source concerns post-alignment large vision-language models during autoregressive decoding, with safety decisions made at the token level. Evidence comes from multiple safety and general-purpose benchmarks; deployment conditions and broader guarantees are not specified.research paper · Sep 3, 2026
Related questions
How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can we uncover unsafe physical behaviors in vision-language-action models before deployment?How can we predict and interpret vision-language model failures to support timely human intervention?How should vision-language models answer valid parts of compound queries while withholding unsafe or unanswerable parts?