arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.05034cs.IR

智能体检索增强生成评估:问题、轨迹与读取之间的预算分配

Agentic RAG Evaluation: Budget Allocation Across Questions, Trajectories, and Reads

Jingjie Ning, Xueqi Li, Yibo Kong

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过HotpotQA和MuSiQue上的检索反馈比较,评估智能体RAG中问题、轨迹和读取的预算分配,发现更广泛的问题覆盖能显著降低标准误差,并提出了基于预测的预算分配方法。

中文摘要 AI 辅助

智能体检索增强生成中的评估预算涵盖问题、搜索轨迹和重复答案。我们使用HotpotQA和MuSiQue上的检索反馈比较来衡量分配精度、读取效率和成本边界。在34.14–34.39M模型令牌下,更广泛的问题覆盖将标准误差降低了33%(相对于五次读取)和12.6%(相对于三条轨迹)。存档的嵌套和仅问题预测分别在这些分配的4.0%和3.5%以内预测了这些分配。深度子集未建立超出两条轨迹审计的明确预测优势。相对于同一令牌预算下的拟合最优值,单次读取的方差惩罚为0–9.9%,且Pro不确定性显著。在记录模型费用下,以每1000次请求0–1美元的搜索价格,更多问题优于更多轨迹;问题与读取的费用排名仍未解决。温度零将答案分歧从14.3%降至3.4%,而比较精度保持相似。关键词:智能体检索增强生成;评估预算;概化理论;重复采样。

英文摘要

Evaluation budgets in agentic retrieval-augmented generation span questions, search trajectories, and repeated answers. We measure allocation precision, reading efficiency, and cost boundaries using a retrieval-feedback comparison on HotpotQA and MuSiQue. At 34.14--34.39M model tokens, broader question coverage lowers standard error by 33\% versus five reads and 12.6\% versus three trajectories. Archived nested and Q-only forecasts predict these allocations within 4.0\% and 3.5\%, respectively. Depth subsets establish no clear forecasting advantage beyond the two-trajectory audit. One-read variance penalties relative to the fitted optimum at the same token budget are 0--9.9\%, with substantial Pro uncertainty. Under recorded model fees, more questions beat more trajectories at search prices of \$0--1 per 1,000 requests; question-versus-read fee rankings remain unresolved. Temperature zero cuts answer disagreement from 14.3\% to 3.4\% while comparison precision stays similar. \par\medskip\noindent\textbf{Keywords:} Agentic RAG; Evaluation budget; Generalizability theory; Repeated sampling.

发表机构

  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

↑