Get Started
Home
Topics
Search
Library
Research questionHow can we interpret multimodal video model performance when benchmark label reliability is unknown?Many video benchmarks publish labels without reliability statistics, making it unclear how much observed error reflects the model versus annotation noise. Dense classroom coding further requires interpreting many behavior labels over time, including codes that may be only partly observable from video.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Machine Learning
Multimodal Models
Latest papersRecent research connected to this question, newest first.VISTA: Dense Multi-Label Classroom Coding with Vision-Language ModelsThe case study uses COPUS to label undergraduate STEM lectures with 24 binary codes every two minutes across 50–90 minutes. Its reference annotations come from a five-person evaluator panel, and the reported evaluation covers three held-out chemistry lectures, with residual errors involving visually similar instructor codes and rare audio-dependent codes.research paper · Sep 3, 2026
Related questions
How can video-language models capture the distribution of human interpretations of dynamic facial expressions?How can multimodal models reason about fine-grained interpersonal relationships from conversational and visual cues?How can long-horizon interactive video world-model training remain reproducible across heterogeneous datasets and incompatible backbones?How can multimodal models maintain useful visual memory for causal streaming video reasoning under fixed memory?