Get Started
Home
Topics
Search
Library
Research questionHow should LLM safety be assessed when jailbreak vulnerability varies by language and persuasive phrasing?English-only tests can overlook safety failures that emerge when harmful requests are expressed in other languages or phrased persuasively. Vulnerability can also differ by harmful-content category, so overall safety scores may conceal specific weaknesses.
AI
Alignment & Safety
Evaluation & Benchmarks
Natural Language Processing
Latest papersRecent research connected to this question, newest first.IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak AttacksThe evidence comes from black-box evaluations of several open-source LLMs using 7,200 adversarial prompts across Hindi, Bengali, Marathi, and Punjabi. The prompts cover ten safety-critical content categories and six human-like persuasive strategies; the findings establish variation across the evaluated languages, strategies, and categories, but do not establish behavior for untested models or languages.research paper · Sep 4, 2026Door-in-the-Face Requests and Refusal Behaviour in Large Language ModelsEvidence comes from nine production models across Anthropic, OpenAI, and Google. Each model received a large request followed by a smaller related request, compared with asking for the smaller request directly; unrelated-topic controls and 265 rewrites from usable instructions to explanations were also tested. The findings do not establish transfer to public benchmark refusals or untested models.research paper · Sep 2, 2026
Related questions
How should LLM safety be evaluated when harmful prompts vary in implicitness and sophistication?How can LLM sandbox security remain reliable when linguistic monitoring misrepresents internal computation?How can LLMs maintain age-appropriate safety for children ages 7–11 across languages and multi-turn conversations?How can black-box LLMs resist jailbreaks without weight access or retraining while preserving benign-query utility?