发表机构
State Key Laboratory for Novel Software Technology, Nanjing University; National University of Singapore; Nanjing University of Posts and Telecommunications(南京大学计算机软件新技术国家重点实验室; 新加坡国立大学; 南京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对长上下文嵌入中均值池化稀释语义信息的问题,提出免训练框架SCSP,通过语义压缩提示选择重要令牌并聚合中间层表示,可即插即用提升零样本与微调模型性能。
AI 中文摘要
大型语言模型(LLMs)在作为免训练文本编码器用于长上下文嵌入方面展现出强大潜力。现有方法主要改进因果注意力下的信息流动,通常通过均匀平均所有令牌表示来构建嵌入。然而,对于长文档,这种均值池化可能会因大量冗余或信息量弱的内容而稀释显著语义信息。为此,我们提出SCSP,一个免训练框架,利用语义压缩在长上下文嵌入中进行信息性令牌选择。具体而言,SCSP首先将文档划分为句子感知的块,并为每个块附加语义压缩提示。提示隔离注意力掩码在保持文档令牌间信息流动的同时,限制每个提示仅关注其对应的局部上下文。然后,我们利用这些提示引发的注意力模式来估计令牌重要性,选择信息性令牌,并将其中间层表示聚合为最终嵌入。在长上下文嵌入基准上的大量实验表明,SCSP可以以即插即用的方式集成到零样本和微调模型中,持续提升其性能。
英文摘要
Large language models (LLMs) have shown strong potential as training-free text encoders for long-context embeddings. Existing approaches primarily improve information flow under causal attention and typically construct embeddings by uniformly averaging all token representations. However, for long documents, such mean pooling can dilute salient semantic information with abundant redundant or weakly informative content. To this end, we propose SCSP, a training-free framework that leverages semantic compression for informative token selection in long-context embedding. Specifically, SCSP first partitions a document into sentence-aware chunks and appends a semantic compression prompt to each chunk. A prompt-isolated attention mask preserves information flow among document tokens while restricting each prompt to its corresponding local context. We then use the attention patterns elicited by these prompts to estimate token importance, select informative tokens, and aggregate their intermediate-layer representations into the final embedding. Extensive experiments on long-context embedding benchmarks demonstrate that SCSP can be integrated into both zero-shot and fine-tuned models in a plug-and-play manner, consistently improving their performance.