Get Started
Home
Topics
Search
Library
Research questionCan user feedback reliably guide LLM revisions if LLM judges overlook the resulting improvements?User interactions may reveal issues that an LLM cannot detect on its own, but the feedback can be noisy and improvements may be difficult to measure. Evaluation becomes especially problematic when judges prefer a baseline response even after feedback has corrected the targeted issue.
AI
Evaluation & Benchmarks
LLM Pretraining & Post-training
Machine Learning
Natural Language Processing
Latest papersRecent research connected to this question, newest first.User Feedback Provides a Unique Signal that LLMs Can not DetectThe evidence compares revisions made with and without feedback on synthetic data with definitive ground truth and on naturalistic data. It also examines cases where LLM judges fail to identify corrections attributable specifically to user feedback.research paper · Sep 2, 2026
Related questions
How can LLMs assess academic proposals and peer feedback with pedagogically grounded scores?How can LLMs generate reliable, adaptive tests that expose one another’s model-specific weaknesses?How can we tell whether agreement among LLM judges reflects human alignment or shared blind spots?How can we reliably assess whether conversational LLMs clarify ambiguity and recover user intent?