arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Shapley上下文剪枝:一种用于上下文重排和剪枝的合作博弈视角

Shapley Context Pruning: A Cooperative Game Perspective for Context Reranking and Pruning

Yanqiao Chen, Dongsheng Hou, Yuhan Rui, Zhen Cao, Yepang Liu

arXiv 2607.16209首次发表:更新:

发表机构

Southern University of Science and Technology(南方科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对检索增强生成系统上下文重排和剪枝缺乏可解释统一框架的问题,提出Shapley上下文剪枝框架,用合作博弈视角归因重要性,采用深度集架构和蒙特卡罗采样,经多实验验证,模型在下游问答性能上表现出色。

AI 中文摘要

上下文重排和剪枝对于提高现代检索增强生成(RAG)系统的效率至关重要,但一个可解释且统一的框架仍未得到充分探索。以往工作主要强调词汇检索、交叉编码器架构、模型蒸馏和低秩适应(LoRA),大多依赖启发式损失函数和经验归因。本文提出了Shapley上下文剪枝(SCP),这是一个用于上下文重排的新框架,通过将上下文建模为合作博弈来建立重要性归因的合作博弈理论视角。在平衡细粒度和粗粒度表示之间的权衡时,我们采用深度集架构在句子级别近似排列不变值函数,利用预训练语言模型作为句子嵌入器并通过成对边际排序损失进行优化。为确保在不牺牲数学严谨性的情况下实现实际可扩展性,我们利用蒙特卡罗采样进行高效训练和推理,为保留Top-K子集排名提供形式化理论误差界和样本复杂度保证。此外,我们进行了全面实验,包括支持句召回、大海捞针(NIAH)评估、长上下文问答和多跳推理,以及对嵌入质量和归因策略的严格消融研究。该模型在与强大基线的对比中实现了有竞争力的下游问答性能。

英文摘要

Context reranking and pruning have become essential for improving the efficiency of modern Retrieval-Augmented Generation (RAG) systems, yet an interpretable and unified framework remains underexplored. Previous work has primarily emphasized lexical retrieval, cross-encoder architectures, model distillation, and Low-Rank Adaptation (LoRA), mostly relying on heuristic loss functions and empirical attribution. This paper presents Shapley Context Pruning (SCP), a novel framework for context reranking that establishes a cooperative-game-theory perspective for importance attribution by modeling the context as a cooperative game. Balancing the trade-off between fine-grained and coarse-grained representations, we employ a Deep Sets architecture to approximate a permutation-invariant value function at the sentence level, utilizing pre-trained language models as sentence embedders and optimizing via a pairwise margin ranking loss. To ensure practical scalability without sacrificing mathematical rigor, we leverage Monte-Carlo sampling for efficient training and inference, providing formal theoretical error bounds and sample complexity guarantees for preserving Top-K subset rankings. Furthermore, we conduct comprehensive experiments-spanning supporting-sentence recall, Needle-in-the-Haystack (NIAH) evaluations, long-context QA, and multi-hop reasoning-alongside rigorous ablation studies on embedding quality and attribution strategies. The model achieves competitive downstream QA performance against robust baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑