Get Started
Home
Topics
Search
Library
Research questionHow can smaller language-model agents resist emotional manipulation while pursuing users’ objectives in adversarial negotiation?Emotionally framed language can shift an agent’s bargaining decisions away from the user’s goals, particularly when the agent is trained to be helpful and accommodating. The difficulty is maintaining effective negotiation behavior while handling emotional tactics as part of an adversarial interaction.
AI
AI Agents
Alignment & Safety
LLM Pretraining & Post-training
Multi-agent Systems
Natural Language Processing
Reinforcement Learning
Small / On-device Models
Latest papersRecent research connected to this question, newest first.EmoDistill: Offline Emotion Skill Distillation for Language Model Agents in Adversarial NegotiationThe source studies smaller language-model agents in adversarial bargaining using offline LLM–LLM interaction data. Evidence covers four emotion-sensitive negotiation domains, a 7B policy, comparisons with vanilla and IQL-only baselines, emotion-free ablations, and transfer to unseen LLM counterparties.research paper · Sep 3, 2026
Related questions
How can language models detect subtle mental manipulation in conversations without missing covert tactics?How can language agents adapt textual world models to evolving behavior in interactive environments without external rewards?Do more capable language models exhibit stable task preferences that conflict with helpful, honest behavior?How can black-box systems detect and mitigate reward hacking in self-evolving language-model loops?