Research questionHow can terminal-agent training environments stay challenging as models improve without costly on-policy synthesis?Scratch-generated environments can provide weak learning signals once agents become capable of solving them. On-policy synthesis requires repeated rollouts and may not generalize well beyond the observed training situations.