Research questionHow can we localize and describe multiple events in untrimmed videos from only ordered event captions?Ordered captions indicate the sequence of events but not when each event occurs, while transitions between events may vary in timing and duration. This makes it difficult to align each description with the correct video segment.