Research questionHow can learned world models prevent policies from exploiting prediction errors and failing in the real world?RL policies can exploit small prediction errors in learned simulators, producing strong simulated performance that fails to transfer to reality. The difficulty is ensuring simulator accuracy where policy decisions make errors consequential.