Research questionHow can on-policy reasoning models use dense self-guidance without reinforcing incorrect solutions?Terminal verifiers give dependable outcome signals but only at the end of a trajectory, whereas dense same-model guidance can amplify false confidence. This tension can cause training to favor incorrect responses, collapse response lengths, or concentrate learning on a narrow set of reasoning strategies.