Get Started
Home
Topics
Search
Library
Research questionHow can multimodal models reason about fine-grained interpersonal relationships from conversational and visual cues?Interpersonal relationships are expressed through subtle social signals distributed across what people say, how they appear, and how they interact. Existing multimodal models lack established ways to measure whether they recover fine-grained relationship dimensions or use the relevant visual evidence.
AI
Computer Vision
Evaluation & Benchmarks
Multimodal Models
Natural Language Processing
Reasoning
Latest papersRecent research connected to this question, newest first.PIVOTSBench: Evaluating Fine-Grained Interpersonal Relationship Reasoning in Multimodal Large Language ModelsThe evidence covers proprietary and open-source multimodal language models evaluated on PIVOTS, built from Social-IQ 2.0 and YouTube data. Auxiliary tasks examine visual-cue identification, while analyses consider visual modalities, explicit social-role information, and joint versus pairwise prediction settings; conclusions are limited to these benchmark tasks and analyses.research paper · Sep 2, 2026
Related questions
How can multimodal models integrate evidence across deeply interleaved text and images?How can multimodal models rely on images or audio rather than language shortcuts?How can multimodal models infer directorial intent from audiovisual choices rather than merely recognize events?How can multimodal agents maintain consistent person identities and reason about relationships across long video memories?