认知分块用于软提示:通过块级因果掩码加速压缩器学习
Cognitive Chunking for Soft Prompts: Accelerating Compressor Learning via Block-wise Causal Masking
- National University of Defense Technology(国防科技大学)
- The Hong Kong Polytechnic University(香港理工大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出PIC方法,通过块级因果掩码加速压缩器学习,提升压缩效率和性能。
AI中文摘要:
通过提示提供大量上下文对于利用大型语言模型(LLMs)的能力至关重要。然而,长上下文显著增加了推理延迟,因为自注意力的计算成本随着序列长度呈平方增长。为缓解这一问题,上下文压缩——特别是软提示压缩——已成为广泛研究的解决方案,它通过训练的压缩器将长上下文转换为较短的记忆嵌入。现有方法通常将整个上下文 indiscriminately 压缩成一组记忆标记,要求压缩器捕捉全局依赖关系,并需要大量预训练数据来学习有效的模式。受人类工作记忆中分块机制的启发,并基于对记忆嵌入相对于原始标记的空间特化的经验观察,我们提出了并行迭代压缩(PIC)。通过简单地修改Transformer的注意力掩码,PIC明确限制记忆标记的接受域到顺序局部块,从而降低压缩器训练的难度。在多个下游任务上的实验表明,PIC在多个下游任务上 consistently 超过竞争基线,其优越性在高压缩场景中尤为明显(例如,在64×压缩比下,PIC在问答任务中实现F1分数相对提升29.8%,EM分数相对提升40.7%)。此外,PIC显著加快了训练过程。具体来说,当训练16×压缩器时,它在超越竞争基线峰值性能的同时,有效将训练时间减少了约40%。
英文摘要:
Providing extensive context via prompting is vital for leveraging the capabilities of Large Language Models (LLMs). However, lengthy contexts significantly increase inference latency, as the computational cost of self-attention grows quadratically with sequence length. To mitigate this issue, context compression-particularly soft prompt compressio-has emerged as a widely studied solution, which converts long contexts into shorter memory embeddings via a trained compressor. Existing methods typically compress the entire context indiscriminately into a set of memory tokens, requiring the compressor to capture global dependencies and necessitating extensive pre-training data to learn effective patterns. Inspired by the chunking mechanism in human working memory and empirical observations of the spatial specialization of memory embeddings relative to original tokens, we propose Parallelized Iterative Compression (PIC). By simply modifying the Transformer's attention mask, PIC explicitly restricts the receptive field of memory tokens to sequential local chunks, thereby lowering the difficulty of compressor training. Experiments across multiple downstream tasks demonstrate that PIC consistently outperforms competitive baselines, with superiority being particularly pronounced in high compression scenarios (e.g., achieving relative improvements of 29.8\% in F1 score and 40.7\% in EM score on QA tasks at the $64\times$ compression ratio). Furthermore, PIC significantly expedites the training process. Specifically, when training the 16$\times$ compressor, it surpasses the peak performance of the competitive baseline while effectively reducing the training time by approximately 40\%.