Research questionHow can reasoning models keep improving on open-ended agentic tasks as human supervision and reliable rewards recede?Mathematics and code often provide automatically verifiable outcomes, but open-ended agentic tasks lack comparably reliable rewards. As human oversight and curated experience become scarce, autonomous feedback and self-generated experience can introduce reward hacking, feedback drift, curriculum collapse, and environment errors.