Get Started
Home
Topics
Search
Library
Research questionHow can black-box systems detect and mitigate reward hacking in self-evolving language-model loops?A self-evolving loop can improve its visible score while the intended capability stagnates or deteriorates. Monitoring must therefore identify proxy exploitation and support corrective update selection without relying on internal model signals.
AI
Alignment & Safety
Evaluation & Benchmarks
Machine Learning
Neural and Evolutionary Computing
Latest papersRecent research connected to this question, newest first.Harness-agnostic detection and immunization of reward hacking in self-evolving language modelsThe source presents a black-box monitor and risk-aware candidate-reselection procedure that operate without access to weights or activations. Evidence comes from a controlled prompt-level host with injected hacking channels and ground-truth labels, so its conclusions are limited to that evaluation setting.research paper · Sep 4, 2026
Related questions
How can reward shaping reduce reward hacking in RLHF when rewards imperfectly capture human preferences?Can black-box attackers identify and reconstruct prompts supposedly removed from language models without knowing them in advance?How can black-box LLMs resist jailbreaks without weight access or retraining while preserving benign-query utility?How can safety-tuned language models distinguish harmful requests from benign ones with risky wording?