arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

BELIEFRAG:在演化证据下使自适应RAG具备状态感知能力

BELIEFRAG: Making Adaptive RAG State-Aware under Evolving Evidence

Hongji Pu

arXiv 2609.39139首次发表:更新:

发表机构

University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对自适应RAG在演化证据下的证据状态碎片化问题,提出BELIEFRAG闭环控制器,显式维护多维度信念状态并决策动作,在多个QA基准上以更少token取得更优F1。

AI 中文摘要

自适应RAG使用置信度、相关性、支持度和检索质量等信号来决定何时检索或修正证据。然而,在多步检索中,这些局部信号必须被整合为对当前证据所支持内容、仍缺失内容以及应采取的后续行动的持续视图。现有方法通常将这些信号用作独立的触发器,难以在轨迹中保持连贯的证据状态;我们将此问题称为证据状态碎片化。我们引入了BELIEFRAG,一种闭环控制器,它更新关于充分性、可靠性、冲突、不确定性、证据缺口和获取成本的显式状态,然后在检索、查询重写、验证、回答、停止和弃权(不执行)之间进行选择。在六个QA基准测试中,使用gpt-oss-120b,BELIEFRAG达到了平均token F1 0.572,每个问题消耗3.89k个token,优于固定迭代检索(F1 0.555),同时少使用39%的token。相同的质量-成本模式也适用于Qwen3-32B,其中BELIEFRAG达到F1 0.552,而迭代检索为0.523,同时少使用35%的token。分析表明,主要收益来自纠正性重新检索而非仅靠剪枝,而多个信念维度是冗余的,校准的可回答性起着最强的操作作用。校准提高了相关证据源之间的阈值稳定性,尽管源偏移仍可能使相同的决策信号失效。

英文摘要

Adaptive RAG uses signals such as confidence, relevance, support, and retrieval quality to decide when to search or correct evidence. In multi-step retrieval, however, these local signals must be combined into a persistent view of what the current evidence supports, what remains missing, and which action should follow. Existing methods often use such signals as separate triggers, making it difficult to preserve a coherent evidence state across a trajectory; we call this problem evidence-state fragmentation. We introduce BELIEFRAG, a closed-loop controller that updates an explicit state over sufficiency, reliability, conflict, uncertainty, evidence gaps, and acquisition cost, then chooses among retrieval, query rewriting, verification, answering, stopping, and abstention. Across six QA benchmarks with gpt-oss-120b, BELIEFRAG reaches mean token F1 0.572 with 3.89k tokens per question, outperforming fixed iterative retrieval (0.555 F1) while using 39% fewer tokens. The same quality-cost pattern transfers to Qwen3-32B, where BELIEFRAG reaches 0.552 F1 versus 0.523 for iterative retrieval while using 35% fewer tokens. Analysis shows that the main gains come from corrective re-retrieval rather than pruning alone, while several belief dimensions are redundant and calibrated answerability plays the strongest operational role. Calibration improves threshold stability across related evidence sources, although source shift can still invalidate the same decision signal.

Comments20 pages, 7 figures, 10 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑