Research questionCan user feedback reliably guide LLM revisions if LLM judges overlook the resulting improvements?User interactions may reveal issues that an LLM cannot detect on its own, but the feedback can be noisy and improvements may be difficult to measure. Evaluation becomes especially problematic when judges prefer a baseline response even after feedback has corrected the targeted issue.