Research questionHow should generated interactive videos be evaluated for action adherence and visual-temporal coherence?Interactive video models must make commanded actions produce the intended scene changes while maintaining coherent appearance and motion. Long videos can obscure the local evidence needed to determine whether individual actions were executed correctly.