arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PC-SubMax:通过正则化子模最大化实现高效提示压缩

PC-SubMax: Efficient Prompt Compression via Regularized Submodular Maximization

Ziyi Zhang, Shuang Cui, Haotian Zhang, Xiaoyu Wang

arXiv 2609.32474首次发表:更新:

AI 中文总结

PC-SubMax将选择性提示压缩建模为正则化单调子模最大化,提出RGM算法,在背包约束下优化信息覆盖、查询相关性与多样性,实现低开销的高效压缩。

AI 中文摘要

虽然大型语言模型(LLMs)越来越多地部署在长上下文场景中,冗长的提示会增加推理成本和延迟,并加剧“中间丢失”现象。选择性提示压缩提供了一种与模型无关的方法来缓解这些问题。然而,基于固定词元或句子级重要性评分的方法可能忽略内容贡献如何随所选子集变化,从而限制了它们解释句子间冗余的能力。依赖自回归LLM评分的压缩过程也可能引入大量开销。我们提出PC-SubMax,一个理论基础的框架,将选择性提示压缩表述为在背包约束下的正则化单调子模最大化。目标函数为$U(S)-\ell(S)$,其中单调子模效用$U$结合了信息覆盖、查询相关性和对数行列式多样性,非负模惩罚$\ell$捕获词元成本。通过边际收益递减,目标函数评估每个句子相对于所选内容的贡献。为优化此目标,我们开发了正则化贪婪+最大(RGM)算法,该算法确定性返回一个可行集$Q$,满足$U(Q)-\ell(Q)\geq \frac{1}{2}U(O)-\ell(O)$,其中$O$是正则化问题的最优可行解。RGM使用$O(n\kappa)$次值预言查询,其中$n$是候选句子数量,$\kappa$是最大可行子集大小。PC-SubMax使用编码器表示,并在压缩过程中避免自回归LLM评分。在七个不同基准上的实验表明,下游性能具有竞争力且压缩开销低。

英文摘要

While large language models (LLMs) are increasingly deployed in long-context scenarios, lengthy prompts can increase inference costs and latency and exacerbate the ``lost-in-the-middle'' phenomenon. Selective prompt compression offers a model-agnostic approach to alleviating these issues. However, methods based on fixed token- or sentence-level importance scores may overlook how content contributions change with the selected subset, limiting their ability to account for inter-sentence redundancy. Compression procedures that rely on autoregressive LLM scoring can also introduce substantial overhead. We propose PC-SubMax, a theoretically grounded framework that formulates selective prompt compression as regularized monotone submodular maximization under a knapsack constraint. The objective is $U(S)-\ell(S)$, where the monotone submodular utility $U$ combines information coverage, query relevance, and log-determinant diversity, and the non-negative modular penalty $\ell$ captures token cost. Through diminishing marginal returns, the objective evaluates each sentence's contribution relative to the selected content. To optimize this objective, we develop the Regularized Greedy+Max (RGM) algorithm, which deterministically returns a feasible set $Q$ satisfying $U(Q)-\ell(Q)\geq \frac{1}{2}U(O)-\ell(O)$, where $O$ is an optimal feasible solution to the regularized problem. RGM uses $O(nκ)$ value-oracle queries, where $n$ is the number of candidate sentences and $κ$ is the maximum feasible subset size. PC-SubMax uses encoder representations and avoids autoregressive LLM scoring during compression. Experiments across seven diverse benchmarks demonstrate competitive downstream performance with low compression overhead.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑