Get Started
Home
Topics
Search
Library
Research questionHow can we localize and describe multiple events in untrimmed videos from only ordered event captions?Ordered captions indicate the sequence of events but not when each event occurs, while transitions between events may vary in timing and duration. This makes it difficult to align each description with the correct video segment.
AI
Computer Vision
Image & Video Processing
Machine Learning
Multimodal Models
Natural Language Processing
Latest papersRecent research connected to this question, newest first.Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video CaptioningThe setting is weakly supervised dense video captioning for untrimmed videos given only ordered event-level captions. The source studies VLM-based transition discovery and temporal-mask refinement, with experiments on ActivityNet Captions and YouCook2.research paper · Sep 3, 2026
Related questions
How can partially relevant video retrieval locate precise query-relevant moments in untrimmed videos with weak supervision?How can multimodal models jointly learn region captioning and spatial localization without text annotations?How can weakly supervised video anomaly detection adapt to anomalies with varying durations and temporal dynamics?How can scalable synthetic video datasets preserve temporal alignment between actions and resulting scene transitions?