Research questionHow can GRPO avoid reinforcing guessed correct answers in bounded-answer and search tasks?GRPO derives rollout advantages from within-group reward statistics, so a correct guess can receive the same strong learning signal as a correctly reasoned solution. Repeated updates may therefore favor guess-like behavior rather than reliable reasoning.