EntroPrefill:基于Renyi引导的上下文剪枝及其对检索增强生成的条件稳定性保证
EntroPrefill: Renyi-Guided Context Pruning with Conditional Stability Guarantees for Retrieval-Augmented Generation
- SRM Institute of Science and Technology(SRM科学与技术学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出EntroPrefill,一种基于Renyi引导的上下文剪枝方法,通过约束丢弃注意力质量并给出理论保证,在检索增强生成中实现安全剪枝,但实验验证留待后续。
AI中文摘要:
中段预填充剪枝可以减少深层Transformer层处理的序列长度,但仅凭注意力集中度并不能证明被丢弃的上下文是可有可无的。我们将EntroPrefill构建为一种Renyi引导的提议机制,并附加对丢弃注意力质量的显式约束。孤立奇异点、正则化的头部池化在尊重分组查询注意力的同时,揭示了专业化与最差头部覆盖之间的定量权衡。我们推导出混合到头部删除包络,一个关于可行令牌移除的可计算上界,以及一个有限样本观测者保证,该保证在自适应选择剪枝层时仍然有效。随后,我们建立了一个条件Transformer扰动界,包含显式的充分Lipschitz常数和一个首令牌决策边际推论。一个反例表明,仅凭浅层观测无法推断无条件未来输出的保证。系统分析区分了查询头联合、物理页面分配和KV传输负载,并给出了剪枝的算术盈亏平衡条件。本文在理论层面展开:定义了流程、假设及其形式限制,但未报告实测加速或任务精度保持。实验留待后续验证假设、近似紧密度及端到端资源权衡。
英文摘要:
Mid-prefill pruning can reduce the sequence processed by deeper transformer layers, but attention concentration alone does not certify that discarded context is dispensable. We formulate EntroPrefill as a Renyi-guided proposal mechanism coupled to explicit constraints on discarded attention mass. Sink-isolated, regularized head pooling respects grouped-query attention while exposing a quantitative trade-off between specialization and worst-head coverage. We derive a mixture-to-head deletion envelope, a computable upper bound on feasible token removal, and a finite-sample observer guarantee that remains valid when the pruning layer is selected adaptively. We then establish a conditional transformer perturbation bound with explicit sufficient Lipschitz constants and a first-token decision-margin corollary. A counterexample shows why shallow observations alone cannot imply an unconditional future-output guarantee. The systems analysis distinguishes query-head unions, physical page allocation, and KV-transfer payload, and gives an arithmetic break-even condition for pruning. This manuscript is theoretical in scope: it defines the procedure, its assumptions, and its formal limits, but does not report measured acceleration or task-accuracy preservation. Experiments are reserved for subsequent validation of the assumptions, approximation tightness, and end-to-end resource trade-offs.