Research questionHow can scalable synthetic video datasets preserve temporal alignment between actions and resulting scene transitions?Action-conditioned video models need each control signal to correspond to the visual transition it causes, but ordinary video rarely records that relationship. Synthetic production must therefore coordinate simulated actions, scene evolution, and rendered observations while scaling across many environments and assets. Latest papersRecent research connected to this question, newest first.Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video GenerationThe evidence concerns an Unreal Engine pipeline for multi-view, action-conditioned video. It separates real-time physics and trajectory recording from offline rendering, and describes distributed production, scene and perceptual-quality filtering, recovery from partial outputs, and cluster monitoring. The reported production used 2,384 asset packs, retained 429 levels and 40 humanoid characters, and generated 2,691 hours of 1080p video plus 6,076 hours of 720p video. The source also notes limitations of perceptual quality proxies for curating world-model data.research paper · Sep 3, 2026