AI 中文总结
提出RAGP方法,将提示压缩建模为多重图上的冗余感知图剪枝,利用莱维游走高效识别非冗余节点,在4倍压缩比下平均得分49.3,优于现有方法。
AI 中文摘要
现有的提示压缩方法将文本视为平坦的标记序列,未能捕捉重要信息的分布式特性,这些信息通常分布在多个位置,并通过局部句法依赖和全局语义关系连接。这种关系结构自然表示为图,其中标记或句子成为节点,其依赖关系成为边。为此,我们提出RAGP,将提示压缩公式化为多重图上的冗余感知图剪枝,该多重图联合建模细粒度注意力依赖和粗粒度语义关系。为了在这种异质结构(密集局部子图和稀疏全局连接)中高效识别非冗余节点,我们采用莱维游走,其重尾步长分布自然地平衡了局部利用与全局探索。在LongBench上的实验表明,RAGP在4倍压缩比下平均得分为49.3,优于现有的基于LLM的压缩方法,例如LongLLMLingua在3倍压缩比下达到48.8。此外,RAGP在多个任务上也超越了最先进的基于视觉的文本压缩范式。代码可在该https URL获取。
英文摘要
Existing prompt compression methods treat text as flat token sequences, failing to capture the distributed nature of important information, which is often spread across multiple locations and connected through both local syntactic dependencies and global semantic relations. Such relational structure is naturally represented as a graph, where tokens or sentences become nodes and their dependencies become edges. To this end, we propose RAGP, which formulates prompt compression as Redundancy-Aware Graph Pruning on a multiplex graph that jointly models fine-grained attention-based dependencies and coarse-grained semantic relations. To efficiently identify non-redundant nodes in this heterogeneous structure (dense local subgraphs and sparse global connections), we employ Levy walks whose heavy-tailed step distribution naturally balances local exploitation with global exploration. Experiments on LongBench show that RAGP achieves an average score of 49.3 under a 4x compression ratio, outperforming existing LLM-based compression methods, such as LongLLMLingua, which attains 48.8 at a 3x compression ratio. Besides, RAGP also surpasses state-of-the-art vision-based text compression paradigms on multiple tasks.