arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24010cs.LG

主动检索生成式人工智能何时应进行检索?效用、校准和成本的预算感知评估

When Should Active RAG Retrieve? A Budget-Aware Evaluation of Utility, Calibration, and Cost

Pin Qian, Su Wang, Chong Peng, Junxian You, Lifei Liu, Haoran Yu, Yihang Chen, Xiaochong Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

研究主动检索生成式人工智能何时检索,通过将其重述为效用估计进行预算感知评估,并分离出相关三个问题,利用多种方法实现,在多数据集和模型中验证,强调评估应报告多方面指标。

中文摘要 AI 辅助

主动检索生成式人工智能(Active RAG)系统在生成过程中决定何时检索外部知识,这使其成为智能检索生成式人工智能(agentic RAG)和自适应检索的预算敏感情况。然而,评估往往未明确操作点:两个系统可能都声称有50%的证据使用预算,但实际的保留使用率不同,因此更高的准确率可能反映的是更宽松的预算,而不是更好的检索策略。我们通过将主动检索重新表述为效用估计来研究主动检索生成式人工智能的预算感知评估,即检索仅通过其相对于无检索答案的边际正确性变化才有价值。这种观点将单点评估混淆的三个问题分开:触发分数是否对有用的检索决策进行排序,根据过去数据校准的阈值是否符合未来预算,以及触发端计算如何改变部署成本。我们通过精确的前k效用前沿、可部署阈值前沿、保守预算前沿、危害审计和成本分解来实现这些问题。在知识密集型多跳问答数据集和开放指令模型中,检索危害不可忽视,路由器排名因数据集和预算而异,名义阈值可能错过目标使用率,简单的不确定性或检索分数基线通常与学习到的效用路由器相当。因此,预算感知的主动检索生成式人工智能评估应在报告准确率的同时,报告前沿、实际使用率、阈值转移误差、危害率和成本分解。

英文摘要

Active RAG systems decide when to retrieve external knowledge during generation, making them a budget-sensitive case of agentic RAG and self-adaptive retrieval. Yet evaluations often leave the operating point underspecified: two systems may both claim a 50% evidence-usage budget while realizing different held-out usage rates, so higher accuracy can reflect a looser budget rather than a better retrieval policy. We study budget-aware evaluation for Active RAG by recasting active retrieval as utility estimation, where retrieval is valuable only through its marginal correctness change over a no-retrieval answer. This view separates three questions that single-point evaluations conflate: whether trigger scores rank useful retrieval decisions, whether thresholds calibrated on past data meet future budgets, and how trigger-side computation changes deployment cost. We operationalize these questions with exact top-k utility frontiers, deployable threshold frontiers, conservative budget frontiers, harm audits, and cost decompositions. Across knowledge-intensive multi-hop QA datasets and open instruction models, retrieval harm is non-negligible, router rankings change across datasets and budgets, nominal thresholds can miss target usage, and simple uncertainty or retrieval-score baselines often rival learned utility routers. Budget-aware Active RAG evaluations should therefore report frontiers, realized usage, threshold-transfer error, harm rates, and cost decompositions alongside accuracy.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)
  • University of Glasgow(格拉斯哥大学)
  • Georgia Institute of Technology(佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑