Get Started
Home
Topics
Search
Library
Research questionHow can we assess vision-language models’ instruction following across languages and instruction-hijacking attacks?Many vision-language evaluations provide limited language coverage and omit adversarial instruction-hijacking scenarios. As a result, model reliability in multilingual, safety-sensitive interactions is difficult to characterize.
AI
Alignment & Safety
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.MM-IFEval-Pro: A Multilingual and Attack-Resistant Benchmark for Instruction-Following in Vision-Language ModelsThe source describes MM-IFEval-Pro, a benchmark with four major task categories, 24 task subcategories, eight instruction categories, 52 instruction subcategories, and an average of three constraints per sample. It covers Chinese and English instructions and diverse instruction-hijacking cases; the source also reports a related reinforcement-learning training set and transfer to other multimodal benchmarks.research paper · Sep 7, 2026
Related questions
How can vision-language models resist multimodal jailbreaks that adapt their strategies and transfer across defenses?How can language models reliably follow instructions containing many simultaneous constraints?How can vision-language models correct unsafe generations token by token without disrupting safe reasoning?How can we predict and interpret vision-language model failures to support timely human intervention?