Research questionHow can reward shaping reduce reward hacking in RLHF when rewards imperfectly capture human preferences?In RLHF, a policy may optimize quirks in a learned reward rather than the behavior humans intend. Reward-shaping choices influence how reliably optimization follows preferences and how stable training remains.