发表机构
Artificial Intelligence for Low-Resource Public Health Application (ALPHA) Centre, Slum and Rural Health Initiative; College of Medicine, University of Ibadan; Department of Computer Science, Faculty of Computing, University of Ibadan; College of Health Sciences, University of Ilorin(低资源公共卫生应用人工智能中心,贫民窟与农村健康倡议; 伊巴丹大学医学院; 伊巴丹大学计算学院计算机科学系; 伊洛林大学健康科学学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出熵驱动自适应路由器(EDAR),利用预测熵在推理时动态选择RAG或长上下文处理,在保留97.4%准确率的同时减少70.7%的token消耗。
AI 中文摘要
现代大语言模型现已支持超过一百万 token 的上下文窗口,这引发了检索增强生成(RAG)是否仍有必要的问题。纯长上下文(LC)处理成本高昂,且已知会忽略位于长输入中间位置的信息;而纯 RAG 速度快,但受限于检索质量,当检索到的块部分相关或相互矛盾时容易出错。我们提出了熵驱动自适应路由器(EDAR),这是一个在推理时决定是从检索块中回答问题还是将其升级为完整长上下文处理的框架。该决策利用基于 RAG 响应前几个生成 token 的 token 级概率分布计算出的预测熵。熵阈值通过在保留验证集上权衡成本与准确率来选定。实验在 LongBench v2 和 Infinity-Bench 上将 EDAR 与纯 RAG 和纯 LC 基线进行了比较。在包含 2,000 个生成的保留集上,预测熵与幻觉率强相关(Pearson r = 0.85,95% 置信区间 [0.83, 0.87])。在长上下文基准上,EDAR 保留了纯长上下文基线准确率的 97.4%,同时将总 token 消耗减少了 70.7%,仅升级了 18.2% 的传入查询。EDAR 与纯长上下文系统之间的准确率差距在标准样本量下与零无统计学显著差异。预测熵是一种有用的模型内部信号,可用于在 RAG 和长上下文推理之间进行路由,基于阈值的混合系统可以以一小部分成本恢复长上下文模型的大部分准确率。该框架不依赖于特定的检索器或 LC 主干,也不需要除解码过程中通常产生的监督之外的额外监督。
英文摘要
Modern large language models now support context windows of more than one million tokens, which has raised the question of whether retrieval-augmented generation (RAG) is still necessary. Pure long-context (LC) processing is expensive and is known to under-attend to information placed in the middle of long inputs, while pure RAG is fast but bounded by retrieval quality and prone to errors when retrieved chunks are partially relevant or contradictory. We propose the Entropy-Driven Adaptive Router (EDAR), a framework that decides at inference time whether to answer a query from retrieved chunks or to escalate it to full long-context processing. The decision uses the predictive entropy of the token-level probability distribution computed over the first few generated tokens of the RAG response. The entropy threshold is selected on a held-out validation set by sweeping cost against accuracy. Experiments compare EDAR against pure-RAG and pure-LC baselines on LongBench v2 and Infinity-Bench. Predictive entropy correlates strongly with hallucination rate on a held-out set of 2,000 generations (Pearson r = 0.85, 95% CI [0.83, 0.87]). On the long-context benchmarks, EDAR retains 97.4% of the accuracy of the pure long-context baseline while reducing total token expenditure by 70.7%, escalating only 18.2% of incoming queries. The accuracy gap between EDAR and the pure long-context system is not statistically distinguishable from zero at standard sample sizes. Predictive entropy is a useful model-internal signal for routing between RAG and long-context inference, and a threshold-based hybrid system can recover most of the accuracy of long-context models at a small fraction of the cost. The framework does not depend on a specific retriever or LC backbone, and it does not require additional supervision beyond what is normally produced during decoding.
Comments11 pages, 2 figures