Get Started
Home
Topics
Search
Library
Research questionCan language models infer others’ mental states as social interactions evolve under unreliable information?Socially grounded tasks require models to use interaction history, infer what participants know or intend, and distinguish reliable from unreliable information. Existing evaluations often isolate these demands, making performance in changing social environments difficult to characterize.
AI
Evaluation & Benchmarks
Multi-agent Systems
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.SocialMaze: A Benchmark for Evaluating and Enhancing Social Reasoning in Large Language Models in Complex Social EnvironmentsSocialMaze evaluates twelve proprietary and open-weight language models on six tasks spanning social deduction games, daily-life interactions, and digital community platforms. The tasks vary in deep reasoning, dynamic interaction, and information uncertainty, with automated checks and human validation supporting data quality. Reported results cover models’ use of evolving histories, performance under uncertainty, reasoning-workflow effects, and targeted fine-tuning, while transfer to language-aggregation tasks remains statistically inconclusive.research paper · Sep 2, 2026
Related questions
How can multimodal models reason about fine-grained interpersonal relationships from conversational and visual cues?How can language models perform complex logical reasoning without accumulating token-level errors?How can we distinguish decodable logical validity from reasoning that actually drives a language model’s answers?How can language models reason iteratively to diagnose complex clinical cases safely and accurately?