Get Started
Home
Topics
Search
Library
Research questionHow can language models detect subtle mental manipulation in conversations without missing covert tactics?Mental manipulation can appear through subtle, covert conversational tactics that are difficult to distinguish from ordinary interaction. Missed instances can limit the usefulness of NLP safety systems in mental-health contexts.
AI
Alignment & Safety
Health
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Detecting Conversational Mental Manipulation with Intent-Aware PromptingThe source concerns LLM-based detection of mental manipulation in conversations and reports experiments on the MentalManip dataset against other prompting strategies. It also states that the code is available.research paper · Sep 3, 2026
Related questions
How can smaller language-model agents resist emotional manipulation while pursuing users’ objectives in adversarial negotiation?Can language models infer others’ mental states as social interactions evolve under unreliable information?How can mixture-of-experts language models preserve safety when adversaries manipulate sparse expert routing?How can language models answer sensitive prompts helpfully without compromising safety?