arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30087cs.CLcs.IRcs.LG

返回还是修订?学习何时修订有助于检索增强问答

Return or Revise? Learning When Revision Helps Retrieval-Augmented QA

Nicholas Kashani Motlagh, Tim Anderson, Jeremy Gwinnup, Grant Erdmann

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对检索增强问答中的答案修订决策,提出基于配对结果的可恢复性预测策略,在多个设置下优于始终修订并缩小理想差距,但需考虑备选方案的影响。

中文摘要 AI 辅助

我们考虑在答案修订系统中,是返回现有草稿答案还是使用检索到的证据对其进行修订的决策问题。草稿置信度估计当前答案是否正确,但该决策需要估计特定修订的效果。为了进行离线训练和评估,我们在相同的正确性评判标准下对返回的草稿及其候选修订进行评分,这使得修复、损害以及与理想答案之间的差距可观测。我们将这种配对效应称为其可恢复性,并训练策略在修订前预测该效应。在三个修订设置下的25,870个保留的开放域问题上,基于配对结果训练的评分器在准确率-修订率曲线下的面积方面,在全部九个Llama设置-种子拟合中均优于匹配的草稿正确性评分器,并在开发集选择的阈值下平均获得0.23-0.68个准确率点的提升,该差异仅在密集检索的训练运行中显著。由此产生的策略优于始终修订,并且平均缩小了超过三分之一的理想差距,尽管它仍然应用了38%-46%的有害修订。然而,当无草稿的标准RAG答案也可用时,在草稿和该答案之间进行选择对于Llama约强两个点,对于OLMo约强四个点,而将候选修订作为第三个选项并未带来显著增益。可恢复性描述了一次修订;其作为可用行动的价值还取决于其他备选方案。

英文摘要

We consider the decision of whether to return an existing draft answer or revise it using retrieved evidence, as in answer-revision systems. Draft confidence estimates whether the current answer is correct, but the decision requires estimating the effect of a specified revision. For offline training and evaluation, we grade both the returned draft and its candidate revision under the same correctness judge, which makes repair, harm, and the gap to an oracle observable. We call this paired effect its recoverability, and we train policies to predict it before revision. On 25,870 held-out open-domain questions across three revision setups, a scorer trained on the paired outcome has greater area under the accuracy--revision-rate curve than a matched draft-correctness scorer in all nine Llama setup--seed fits, and gains 0.23--0.68 accuracy points on average at development-selected thresholds, a difference significant across training runs only for dense retrieval. The resulting policy improves on always revising and on average closes more than a third of the oracle gap, although it still applies 38--46% of the harmful revisions. When a draft-free standard-RAG answer is also available, however, choosing between the draft and that answer is stronger by about two points for Llama and four for OLMo, and adding candidate revision as a third option yields no significant gain. Recoverability describes one revision; its value as an available action also depends on the alternatives.

发表机构

  • DCS Corp(DCS公司)
  • Air Force Research Laboratory(空军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑