预测智能体检索增强生成中的部分答案质量与效用
Predicting Partial Answer Quality and Utility in Agentic Retrieval-Augmented Generation
浏览论文内容
中文总结 AI 辅助
本研究提出轨迹内探测框架,定义部分答案质量与效用,用于预测智能体RAG中间状态,实现提前停止,减少11%迭代且保留98%答案质量。
中文摘要 AI 辅助
智能体检索增强生成(Agentic Retrieval-Augmented Generation, RAG)已成为多跳问答的一种有前景的范式,其中推理模型迭代地向检索器发出查询,并将新检索到的上下文纳入后续推理步骤。虽然这种迭代过程可以提高最终答案质量,但当前对智能体RAG的评估主要关注端到端结果,对模型在生成过程中答案状态的变化提供的可见性有限。在这项工作中,我们引入了一种轨迹内探测框架来研究智能体RAG中的中间答案状态。具体来说,在每次检索-推理迭代之后,我们强制智能体模型停止推理,并基于其当前状态生成一个中间答案。这使我们能够定义两个迭代级度量:每次迭代的部分答案质量,以及部分效用,即部分答案质量在迭代间的变化。我们在多跳问答基准上的分析表明,部分答案质量通常在自然终止之前趋于平稳,许多后续迭代仅带来微小的可测量改进。因此,我们提出了两个预测任务:部分答案质量预测和部分效用预测,并从迭代内、迭代间和查询-迭代视角研究轨迹衍生信号。实验表明,部分答案质量比部分效用更可预测,监督模型在质量预测上实现了皮尔逊相关系数r高于0.43。最后,使用预测的答案质量和效用进行提前停止,将平均迭代次数减少了约11%,同时保留了自然停止所达到的最终答案质量的约98%。
英文摘要
Agentic Retrieval-Augmented Generation (RAG) has become a promising paradigm for multi-hop question answering, where a reasoning model iteratively issues queries to a retriever and incorporates newly retrieved context into subsequent reasoning steps. While this iterative process can improve final answer quality, current evaluations of agentic RAG largely focus on end-to-end outcomes and provide limited visibility into how a model's answer state changes during generation. In this work, we introduce an in-trajectory probing framework to study intermediate answer states in agentic RAG. Specifically, after each retrieval-reasoning iteration, we force an agentic model to stop reasoning and generate an intermediate answer based on its current state. This allows us to define two iteration-level measures: partial answer quality at each iteration, and partial utility as the change in partial answer quality across iterations. Our analysis across multi-hop QA benchmarks reveals that partial answer quality often plateaus before natural termination, with many later iterations contributing only small measurable improvements. Accordingly, we formulate two prediction tasks, partial answer quality prediction and partial utility prediction, and study trajectory-derived signals from intra-iteration, inter-iteration, and query-iteration perspectives. Experiments show that partial answer quality is more predictable than partial utility, with supervised models achieving Pearson's r above 0.43 for quality prediction. Finally, using predicted answer quality and utility for early stopping reduces average iteration count by about 11% while preserving about 98% of the final answer quality achieved by natural stopping.
发表机构
- University of Glasgow(格拉斯哥大学)
机构由 AI 辅助整理,请以论文原文为准。