Research questionHow can group-relative policy optimization enforce constraints without normalization coupling reward and constraint objectives?In constrained GRPO, normalizing scalarized rewards within each sampled group can make changing one constraint multiplier alter the effective weighting of the reward and other constraints. This coupling can destabilize multiplier dynamics and make constraint adherence harder to maintain while optimizing task performance.