Research questionHow can non-incremental tree learners handle distribution shift in online self-play reinforcement learning?As self-play policies improve, the distribution of observed game states changes continuously. Non-incremental tree learners typically require batch refitting, making it difficult to adapt to this shifting data during reinforcement learning.