AI 中文总结
研究探讨仅用语言规则作提示压缩器的可行性,通过离线进化搜索找到有竞争力的规则组合,所产生的语言压缩器仅用CPU端处理,经双路径协议评估性能与先进策略相似,揭示了压缩时规则变化及性能随压缩力度的表现。
AI 中文摘要
提示压缩可缩短大语言模型(LLM)输入以降低推理成本,但现有方法通过语言模型前向传递来评估令牌重要性,这种细微且成本高的令牌选择是否必要仍存疑问。压缩需要识别信息内容,语言学研究长期以来通过可转化为确定性规则的线索解决此问题。因此研究人员提出疑问:仅语言规则能否作为有效的提示压缩器,在压缩时无需基于语言模型的评分?为解决此问题,研究人员对词汇、句法、语义和语篇种子进行离线进化搜索以找到有竞争力的规则组合。由此产生的语言压缩器在部署时无需语言模型前向传递,仅使用CPU端处理进行压缩。研究人员用双路径协议评估它以平衡压缩质量和重建保真度。在短文、多文档推理和对话记忆问答数据集上,进化后的压缩器性能与近期先进的提示压缩策略相似。性能在轻度到中度压缩下最强,随着压缩力度加大而下降,直接路径和重建路径呈现不同模式。进化分析表明,有效的压缩融合了跨语言层次的信号,且随着压缩率增加,规则从令牌修剪转向句子提取。
英文摘要
Prompt compression shortens LLM input to reduce inference cost, yet existing methods score token importance through LM forward passes. It remains questionable whether such nuanced, costly token selection is necessary. Compression requires identifying informative content, a problem that linguistic research has long addressed through cues that can be operationalized as deterministic rules. We therefore ask: can \textbf{linguistic rules alone} serve as effective prompt compressors, without LM-based scoring at compression time? To address this, we conduct offline evolutionary search over lexical, syntactic, semantic, and discourse seeds to find competitive rule combinations. The resulting linguistic compressor requires no LM forward pass at deployment and uses only CPU-side processing for compression. We evaluate it with a dual-path protocol to balance compression quality and reconstruction fidelity. Across short passages, multi-document reasoning, and dialogue-memory QA datasets, evolved compressors achieve performance similar to that of recent advanced prompt-compression strategies. Performance is strongest under light-to-moderate compression and degrades as compression becomes more aggressive, while the Direct and Reconstruction paths exhibit distinct patterns. Evolutionary analysis reveals that effective compression fuses signals across linguistic levels and, as the compression ratio increases, rules shift from token pruning to sentence extraction.
Comments37 pages, 6 figures, under review paper