大型语言模型的词汇提示压缩:一种免训练、确定性的流水线及跨十一类任务的实证帕累托分析
Lexical Prompt Compression for Large Language Models: A Training-Free, Deterministic Pipeline with Empirical Pareto Analysis Across Eleven Task Categories
浏览论文内容
中文总结 AI 辅助
本研究提出一种免训练、确定性、仅CPU的词汇提示压缩流水线,通过十种可切换的词汇转换在1,242个提示上实现最高40.3%令牌减少,同时保持高保真度,并刻画了跨十一类任务的压缩-保真度帕累托前沿。
中文摘要 AI 辅助
近年来,大型语言模型(LLM)的进步使得提示词变得日益庞大和复杂。链式思维推理(Wei等人,2022)和上下文学习(Brown等人,2020)等技术经常将现实世界的提示词推至数千个令牌,从而增加了推理成本和延迟。诸如LLMLingua(Jiang等人,2023)和选择性上下文(Li等人,2023)等学习型压缩方法实现了高压缩比,但需要辅助语言模型且是非确定性的。我们提出了一个互补的问题:在输出质量显著下降之前,基于经典词汇NLP的免训练、完全确定性、仅CPU的流水线能推进到什么程度?我们组装了一个可配置的流水线,包含十种可切换的词汇转换:停用词移除、填充短语删除、缩写和缩略词替换、基于词性的剪枝、词形还原、WordNet驱动的同义词缩短以及命名实体保留。我们在来自六个来源(Dolly-15k、LMSYS-Chat-1M、WildChat-1M、MMLU、GSM8K、HellaSwag)的1,242个仅英文提示上评估了十五种配置,这些提示跨越了十一个自动派生的任务类别,产生了18,630对GPT-4o-mini补全。输出保留使用BLEU、ROUGE-1/2/L、BERTScore-F1和SentenceBERT余弦相似度进行衡量。最激进的配置在BERTScore-F1为0.876(相对于原始提示输出)的情况下实现了平均令牌减少40.3%(标准差=9.2);仅停用词配置在0.913的BERTScore-F1下实现了29.6%的减少。压缩与保真度之间的帕累托前沿按任务类别进行了刻画,其中常识推理在激进压缩下是系统性失败模式。所有代码、提示和逐单元结果均已发布以供复现。
英文摘要
Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et al., 2022) and in-context learning (Brown et al., 2020) frequently push real-world prompts past several thousand tokens, increasing inference cost and latency. Learned compression methods such as LLMLingua (Jiang et al., 2023) and Selective Context (Li et al., 2023) achieve high compression ratios but require auxiliary language models and are non-deterministic. We ask a complementary question: how far can a training-free, fully deterministic, CPU-only pipeline based on classical lexical NLP be pushed before output quality degrades significantly? Eleven toggleable lexical transformations - stopword removal, filler-phrase deletion, contraction and abbreviation substitution, part-of-speech-based pruning, lemmatization, WordNet-driven synonym shortening, and named-entity preservation - are assembled into a configurable pipeline. Fifteen configurations are evaluated on 1,242 English-only prompts from six sources (Dolly-15k, LMSYS-Chat-1M, WildChat-1M, MMLU, GSM8K, HellaSwag), spanning eleven automatically derived task categories, yielding 18,630 paired GPT-4o-mini completions. Output preservation is measured using BLEU, ROUGE-1/2/L, BERTScore-F1, and SentenceBERT cosine similarity. The most aggressive configuration achieves a mean token reduction of 40.3% (sigma = 9.2) at a BERTScore-F1 of 0.876 against the original-prompt output; a stopword-only configuration achieves 29.6% reduction at 0.913. The compression-versus-fidelity Pareto frontier is characterized per task category, with commonsense reasoning a systematic failure mode under aggressive compression. All code, prompts, and per-cell results are released for reproducibility.
发表机构
- Bright Horizons(光明地平线)
- Northeastern University(东北大学)
机构由 AI 辅助整理,请以论文原文为准。