Get Started
Research questionHow can language models assigned protective roles avoid claiming real-world actions they cannot perform?A model may claim to have contacted emergency services or administered care despite lacking physical or institutional agency. Such claims can mislead vulnerable users about whether protection has actually occurred.
AI
Alignment & Safety
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Protective Capacity Hallucination: When Large Language Models Claim Nonexistent CapabilitiesThe evidence covers eight LLMs and 13,600 sessions across different situational severities and interactional formats. It reports near-ceiling hallucination levels in most models for multi-party dialogue in ordinary service domains, but floor levels in intimate-partner conflict scenarios; it does not establish a particular mitigation or broader deployment performance.research paper · Sep 4, 2026
Related questions
How can safety evaluations measure language-model behavior without triggering evaluation-aware changes in decisions?How can we uncover unsafe physical behaviors in vision-language-action models before deployment?How can LLM-agent systems prevent safety compromises from propagating across workflow boundaries?How can we predict and interpret vision-language model failures to support timely human intervention?
Home
Topics
Search
Library