Get Started
Home
Topics
Search
Library
Research questionHow can video continuation models apply new prompts to world states that past frames do not reveal?Segmented video generators often carry forward frames or features, but these records may not encode the state produced by earlier actions. The next continuation can therefore contradict changes that are occluded or only implied by the preceding video and prompt.
AI
Computer Vision
Evaluation & Benchmarks
Image & Video Processing
Reasoning
Video Generation
Latest papersRecent research connected to this question, newest first.Do Video Generators Track the World Across Segments? A Benchmark and Method for World-State Reasoning in Video ContinuationThe source evaluates continuation from a previous video, its prompt, and a new prompt using the Statebench benchmark, and reports results for a proposed Stateagent system. Evidence covers past-visible, occluded-process, and complex-transition states, with additional results for one-minute story generation.research paper · Sep 3, 2026
Related questions
How can text-promptable video segmentation track targets through disappearance while rejecting visually similar artifacts?How can continual VideoQA learn new tasks without forgetting earlier ones or accumulating task-specific prompts?How can long-horizon interactive video world-model training remain reproducible across heterogeneous datasets and incompatible backbones?How can multimodal world models maintain physically coherent 3D scenes during closed-loop interaction?