发表机构
The Hebrew University of Jerusalem; IBM Research; MIT; MIT-IBM Watson AI Lab(耶路撒冷希伯来大学; IBM研究院; 麻省理工学院; 麻省理工-IBM沃森人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出用户反馈是LLMs可利用的高价值改进信号,其被低估源于评估范式偏差,反馈驱动的模型修订解决问题比率更高,且LLM评判者无法识别反馈带来的修正。
AI 中文摘要
利用用户交互产生的自然反馈为大型语言模型(LLMs)提供了极具潜力的学习信号,但近期研究表明这类反馈存在固有噪声,难以有效利用。我们对这一观点提出挑战,证明用户反馈是可用于改进的高可行性信号,其感知到的无效性源于当前评估范式中的系统性偏差。为分离反馈的实用性,我们构建了带有明确真实值的合成数据,同时使用自然数据验证研究结果适用于现实场景。通过对比有无反馈时生成的模型修订,我们发现反馈驱动的修订解决目标问题的比率显著高于基线修订。最后,我们揭示了评估偏差的根源:当模型仅因反馈成功修复问题时,LLM评判者常无法识别真正修正的响应,反而系统性偏好性能更差的基线输出。
英文摘要
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.