发表机构
Ahsanullah University of Science and Technology(阿赫桑乌拉科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对固定读者上下文预算下的检索增强生成,提出答案在上下文中诊断指标,并构建基于子模最大化的打包器,在特定条件下提升性能。
AI 中文摘要
在固定读者上下文预算下的检索增强生成(RAG)面临一个选择问题:在检索到的证据中,只有一部分可以展示给读者。我们认为文档召回——标准的检索指标——在这个场景下是错误需要优化的量,并做出两个贡献。首先,作为一般贡献,我们引入了答案在上下文中(answer-in-context),一个诊断指标,用于衡量黄金答案是否作为连续片段存在于打包的读者上下文中(而非检索集中)。它比召回率更好地预测答案F1(r=0.39-0.55 vs. 约0.31),将答案质量大致分离五倍(HotpotQA上0.60 vs. 0.12),并携带检索之外的信息:它在召回率之上增加Delta R平方=0.17,并且即使在所有黄金都被检索到的问题中,也显示出4.6倍的EM差距。我们还通过干预性实验确认了这一点:在2WikiMultiHopQA上,一个提高覆盖率但不提高答案在上下文中的打包变化并未带来准确率提升。其次,作为条件贡献,我们将读者上下文构建视为预算约束的单调子模最大化问题,并构建了一个打包器,联合优化相关性、查询覆盖率、代表性和多样性。在HotpotQA上,使用160个token的预算和3B读者,它击败了强聚焦启发式方法、MMR和朴素打包——在相同或更低的token成本下,跨三个种子,F1最多提升+5.1。关键的是,我们诚实地映射了这一胜利的范围:它需要(i)多跳互补结构,(ii)检索出证据,(iii)有约束但非极端的预算,以及(iv)读者足够弱,使得证据密度而非阅读能力成为瓶颈。一个量化控制的读者规模阶梯(3B到7B到14B)显示,相对于启发式方法的优势在7B时被吸收,并在14B时显著逆转,而诊断指标用一个变量解释了每个边界。
英文摘要
Retrieval-augmented generation under a fixed context budget forces a selection problem: only a fraction of the retrieved evidence fits in front of the reader. The field's standard metric, recall@k, is scored on the retrieved set, but the reader consumes the packed context - and once packing must discard evidence, the two come apart. We introduce answer-in-context, a diagnostic that measures whether a gold answer survives into the packed context, and argue it is the quantity budgeted RAG should be optimizing. It carries substantial information beyond retrieval, adding Delta R^2 = 0.17-0.27 over recall across three multi-hop datasets; even among questions where all gold was retrieved, whether packing keeps the answer separates exact match by 4.6x. Two independent interventions confirm the mediation: a packing change that raises document coverage without raising answer-in-context leaves accuracy flat, and prompt compression that destroys the answer span lowers both together. A graded variant extends the diagnostic to free-form answers, where no verbatim span exists. We then show the diagnostic is actionable. Casting reader-context construction as budgeted submodular maximization gives a packer that beats both deployed top-k truncation and LLMLingua-2 compression - across three reader families, four scales, and four budgets, at equal-or-lower token cost. Against a hand-tuned query-focused heuristic, which we show approximates the same objective, it reaches parity, winning outright only where evidence density is the binding constraint. Throughout, one variable predicts what helps and what cannot.
CommentsUnder review at EACL 2027