Research questionHow can black-box systems detect and mitigate reward hacking in self-evolving language-model loops?A self-evolving loop can improve its visible score while the intended capability stagnates or deteriorates. Monitoring must therefore identify proxy exploitation and support corrective update selection without relying on internal model signals.