arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

上下文分配定律:生成式搜索中的因果测量与闭环编排

The Laws of Context Allocation: Causal Measurement and Closed-Loop Orchestration in Generative Search

Peiyang Liu, Xi Wang, Di Liang, Wei Ye

arXiv 2608.23252首次发表:更新:

发表机构

National Engineering Research Center for Software Engineering, Peking University; Peking University; Tencent(北京大学国家软件工程研究中心; 北京大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对RAG的证据利用测量与上下文预算分配瓶颈,提出因果留一法探针与闭环次模调度器,实现组合召回率提升16.7-20.5个百分点,确立顺序反馈驱动编排为生成式搜索的范式。

AI 中文摘要

随着检索增强生成(RAG)向多样化组合生成方向发展,它受到两个关键瓶颈的阻碍:证据利用的有缺陷测量,以及次优的上下文预算分配。我们依次解决这两个问题。为解决测量问题,我们揭示了一种普遍存在的“诊断错觉”:标准相关性代理在困难负例上表现极差。我们用一种高效的因果留一法探针取代它们,该探针能准确分离生成式依赖,并正式校准大语言模型(LLM)注意力的结构稀释。为解决分配问题,我们在去混淆因子网格中部署该因果探针。我们证明,整体上下文拓宽的主流策略是一种受相关性衰减惩罚的架构陷阱。相反,在多个顺序生成中迭代分配计算,可带来16.7至20.5个绝对百分点的组合召回率提升,且在32B模型上稳健扩展。最后,我们将这些解决方案统一为可部署的闭环次模调度器。由归因引导的对比解码器增强以覆盖LLM注意力惯性,我们的架构系统地强制整合新证据。通过超越经典开环基线,我们确立了顺序、反馈驱动的编排作为生成式搜索的决定性范式。我们的代码、数据和因果测量工具可在该httpsURL获取。

英文摘要

As Retrieval-Augmented Generation (RAG) shifts toward diverse portfolio generation, it is stymied by two critical bottlenecks: flawed measurement of evidence utilization, and suboptimal context budget allocation. We resolve both sequentially. To resolve measurement, we expose a pervasive ``diagnostic illusion'': standard relevance proxies fail catastrophically on hard negatives. We replace them with an efficient causal leave-one-out probe that accurately isolates generative reliance and formally calibrates the structural dilution of LLM attention. To resolve allocation, we deploy this causal probe in a deconfounded factorial grid. We prove that the prevailing strategy of monolithic context widening is an architectural trap penalized by relevance decay. Instead, allocating compute iteratively across multiple sequential generations drives transformative portfolio recall gains of 16.7--20.5 absolute percentage points, scaling robustly up to 32B models. Finally, we unify these solutions into a deployable closed-loop submodular scheduler. Augmented by an attribution-steered contrastive decoder to override LLM attention inertia, our architecture systematically forces fresh evidence integration. By dominating classical open-loop baselines, we establish sequential, feedback-driven orchestration as the definitive paradigm for generative search. Our code, data, and causal measurement instruments are available at https://github.com/PeiYangLiu/ascp.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑