证据接口塑造了检索增强型读者使用支持信息的方式
Evidence Interfaces Shape How Retrieval-Augmented Readers Use Support
浏览论文内容
中文总结 AI 辅助
研究多跳 RAG 评估中前 k 答案分数隐藏的问题,通过对比不同训练方式的适配读者,区分支持可用性失败与读者接口效应,并指出在支持标注评估中应同时报告前 k 答案分数与完整支持覆盖率,黄金支持优先等方法可改进读者。
中文摘要 AI 辅助
在多跳检索增强生成(RAG)评估中,前 k 个答案分数可能隐藏两种不同的失败情况:检索窗口可能遗漏部分支持链,或者可能包含适配读者使用不佳的支持形式。我们将这种面向读者的检索证据形式称为证据接口。使用三个带有支持标注的多跳问答基准,我们比较了分别用原始上下文、检索窗口和黄金支持诊断渲染训练的匹配适配读者。这些比较区分了支持可用性失败与其余读者接口效应。只有在检查完整标注的支持链是否存在后,前 k 个窗口才变得可解释:当存在时,短排名窗口可以匹配或优于原始上下文;当不存在时,缺失的支持解释了大部分损失。黄金支持优先改进匹配读者;在 2Wiki 和 MuSiQue 上,支持监督排序器以较低的提示成本提高覆盖率并恢复原始上下文质量,同时保留黄金空间。支持去除检查进一步表明收益依赖于公开的证据,而不仅仅是答案先验。因此,在支持标注的评估中,前 k 个答案分数应与完整支持覆盖率一起报告。
英文摘要
In multi-hop RAG evaluation, a top-k answer score can hide two different failures: the retrieval window may drop part of the support chain, or it may contain support in a form the adapted reader does not use well. We call this reader-facing form of retrieved evidence an evidence interface. Using three support-annotated multi-hop QA benchmarks, we compare matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings. These comparisons distinguish support-availability failures from remaining reader-interface effects. Top-k windows become interpretable only after checking whether the complete annotated support chain survives: when it does, short ranked windows can match or improve over raw context; when it does not, missing support explains much of the loss. Gold support-first improves matched readers; on 2Wiki and MuSiQue, a support-supervised ranker raises coverage and recovers raw-context quality at lower prompt cost, while retaining gold headroom. Support-removal checks further show that the gains rely on exposed evidence, not only answer priors. On support-annotated evaluations, top-k answer scores should therefore be reported together with complete-support coverage.