arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.11238cs.AI

面向查询不可知的RAG评估:基于查询覆盖率与声明可验证性

Towards Query-Agnostic RAG Evaluation via Query Coverage and Claim Verifiability

Jeonghwan Choi, Taewon Yun, Minjeong Ban, Gyeonghun Sun, Jae-Gil Lee, Hwanjun Song

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Q-CARE框架,通过分解查询与答案实现RAG的细粒度评估,在8个数据集的基准测试中,其评估结果与人工判断的相关性优于现有4种RAG评估指标。

中文摘要 AI 辅助

检索增强生成(RAG)通过将模型响应建立在检索到的证据基础上,提升了大语言模型的事实准确性,但现有评估框架难以针对从封闭式事实查询到开放式解释请求的各类用户查询,提供一致、细粒度的诊断。本文提出Q-CARE,这是一个查询不可知且完全无需参考的框架,它通过将查询分解为子查询、将答案分解为原子声明来实现细粒度评估。Q-CARE基于查询覆盖率与声明可验证性确立了统一评估原则,生成了感知覆盖率的检索器指标(C-Prec@k、C-nDCG@k)和声明级生成器指标(完整性、简洁性与可验证性)。在涵盖8个数据集的人工标注基准上,Q-CARE与人工判断的相关性高于包括RAGEval和RAGChecker在内的4种现有RAG评估指标,证明其作为可靠自动化评估框架的有效性。代码与数据可在指定URL公开获取。

英文摘要

Retrieval-augmented generation improves the factuality of large language models by grounding responses in retrieved evidence, yet existing evaluation frameworks struggle to provide consistent, fine-grained diagnostics across the diverse spectrum of user queries, ranging from close-ended fact-seeking to open-ended explanatory requests. We propose Q-CARE, a query-agnostic and fully reference-free framework that enables fine-grained assessment by decomposing queries into sub-queries and answers into atomic claims. Q-CARE establishes a unified evaluation principle based on query coverage and claim verifiability, yielding coverage-aware retriever metrics (C-Prec@k, C-nDCG@k) and claim-level generator metrics (Completeness, Conciseness, and Verifiableness). On a human-annotated benchmark spanning eight datasets, Q-CARE achieves higher correlation with human judgments than four existing RAG evaluation metrics, including RAGEval and RAGChecker, proving its effectiveness as a reliable, automated evaluation framework. Code and data are publicly available at https://github.com/DISL-Lab/Q-CaRE-COLM-26.

补充信息

↑