arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2505.22430cs.CL

RAG-Zeval:通过端到端规则引导推理实现稳健且可解释的 RAG 响应评估

RAG-Zeval: Towards Robust and Interpretable Evaluation on RAG Responses through End-to-End Rule-Guided Reasoning

  • The Chinese University of Hong Kong(香港中文大学)
  • Massachusetts Institute of Technology(麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

Kun Li, Yunxiang Li, Tianhua Zhang, Hongyin Luo, Xixin Wu, James Glass, Helen Meng

更新

AI总结:

本文提出 RAG-Zeval 框架,将 RAG 响应的忠实性与正确性评估构建为规则引导推理任务,利用强化学习与基于排序的结果奖励机制训练紧凑模型,在降低计算成本的同时实现了超越大参数基线的评估效果与卓越可解释性。

AI中文摘要:

稳健的评估对于部署值得信赖的检索增强生成(RAG)系统至关重要。然而,当前基于 LLM 的评估框架主要依赖于直接用复杂的多阶段提示词提示资源密集型模型,未能充分利用模型的推理能力,并引入了显著的计算成本。本文提出了 RAG-Zeval(RAG-Zero Evaluator),一种新颖的端到端框架,将忠实性和正确性评估构建为规则引导的推理任务。我们的方法使用强化学习训练评估器,促使紧凑型模型在一次处理中生成全面且合理的评估及详细解释。我们引入了基于排序的结果奖励机制,使用偏好判断而非绝对分数,以解决获取精确逐点奖励信号的挑战。为此,我们通过生成零人工标注的质量控制响应来合成排序参考。实验表明 RAG-Zeval 表现优异,与人类判断实现了最强的相关性,并优于依赖参数量大 10-100 倍的 LLM 的基线方法。我们的方法在响应评估中还展现出卓越的可解释性。

英文摘要:

Robust evaluation is critical for deploying trustworthy retrieval-augmented generation (RAG) systems. However, current LLM-based evaluation frameworks predominantly rely on directly prompting resource-intensive models with complex multi-stage prompts, underutilizing models' reasoning capabilities and introducing significant computational cost. In this paper, we present RAG-Zeval (RAG-Zero Evaluator), a novel end-to-end framework that formulates faithfulness and correctness evaluation as a rule-guided reasoning task. Our approach trains evaluators with reinforcement learning, facilitating compact models to generate comprehensive and sound assessments with detailed explanation in one-pass. We introduce a ranking-based outcome reward mechanism, using preference judgments rather than absolute scores, to address the challenge of obtaining precise pointwise reward signals. To this end, we synthesize the ranking references by generating quality-controlled responses with zero human annotation. Experiments demonstrate RAG-Zeval's superior performance, achieving the strongest correlation with human judgments and outperforming baselines that rely on LLMs with 10-100 times more parameters. Our approach also exhibits superior interpretability in response evaluation.

↑