Research questionHow can long-horizon interactive video world-model training remain reproducible across heterogeneous datasets and incompatible backbones?Interactive video world models combine datasets with different temporal scales, camera geometries, visual quality, motion, and captioning. Backbone-specific representations and processing pipelines can make supervision inconsistent and results difficult to reproduce or compare. Latest papersRecent research connected to this question, newest first.SolarWM: Open Data and Scalable Training for Long-Horizon Video World ModelsThe source describes an open system spanning data preparation through long-horizon inference, using 1.43 million clips from 10 datasets, shared camera-conditioning, training, and inference interfaces, and four 5B–33B models based on Wan2.2, LTX-2.5, and MiniMax-H3. It reports causal rollouts lasting minutes to hours after training on 5-second sequences and releases the data, pipelines, recipes, weights, and framework; the supplied evidence does not establish broader performance beyond these reported configurations.research paper · Sep 2, 2026