Research questionWhen a model-based agent performs well, how can we verify its visual world model supports reliable planning and transfer?A high return can conceal visual or dynamical errors in the model used to imagine future trajectories. Such errors may become apparent when a new policy learns entirely inside a frozen model and is then evaluated in the real environment.