发表机构
Hostinger; Kaunas University of Technology(Hostinger; 考纳斯理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究审计抽取式提示压缩器的跨语言性能,发现英文训练的压缩器存在跨语言迁移差距,多语言训练的 XProvence 无此差距,先翻译再压缩的 pipeline 可在部分语言中以更低成本实现相当或更优效果。
AI 中文摘要
抽取式提示压缩有望通过移除低信息 token 降低大语言模型(LLM)的推理成本,而 LLMLingua-2 等学习型压缩器在英文基准上表现出色。其他大多数语言已存在 token 溢价:相同内容的 token 消耗是英文的 1.3-1.8 倍。本文探究压缩是缩小还是扩大了这一差距。我们使用覆盖五种书写系统的十种语言的完全平行数据,在目标模型的分词器中控制预算匹配,审计了四种学习型压缩器和四种确定性基线,涉及来自十个厂商的十一个目标模型(超过 25 万次评估调用)。其中三种压缩器使用英文监督训练(LLMLingua-2 XLM-R/mBERT;生产级 Headroom 栈的 Kompress-v2);第四种 XProvence 为多语言训练。研究发现:一是迁移差距真实存在,在目标模型和压缩器主干中均有复现,且与压缩率高度相关:在 0.33 的保留率下,英文保留了 57-62% 的归一化上下文利用率,立陶宛语仅保留 10-24%,中文则基本为零,尽管中文的 token 溢价最小。二是该差距与压缩监督数据相关,而非架构:所有三种英文训练的压缩器均显示此差距,确定性方法无可比差距,多语言训练的 XProvence v1 无此差距;其 v2 版本在翻译数据上重新训练,在激进阈值下清空了 92% 的中文上下文且无任何预警。三是在更困难的长上下文设置中,激进的学习型压缩使五种非英语语言中的三种的压缩上下文效用降至或低于无上下文水平。此外,先翻译再压缩的 pipeline 在五种测试语言中的三种中, token 成本约为原生压缩的一半,效果相当或更优。我们发布了所有代码、压缩结果和模型输出,非英语语言的安全压缩预算要小得多。
英文摘要
Extractive prompt compression promises to cut LLM inference costs by removing low-information tokens, and learned compressors such as LLMLingua-2 report strong results on English benchmarks. Most other languages already pay a token premium: the same content costs 1.3-1.8x more tokens than in English. We ask whether compression closes or widens this gap. Using fully parallel data in ten languages spanning five scripts, with controls budget-matched in the target model's tokenizer, we audit four learned compressors against four deterministic baselines, on eleven target models from ten vendors (over 250,000 evaluation calls). Three of the compressors are trained with English supervision (LLMLingua-2 XLM-R/mBERT; Kompress-v2 from the production Headroom stack); the fourth, XProvence, is trained multilingually. First, the transfer gap is real, replicates across target models and compressor backbones, and is strongly rate-dependent: at a 0.33 keep-rate English retains 57-62% of normalized context utilization while Lithuanian retains 10-24% and Chinese essentially none, despite Chinese having the smallest token premium. Second, the gap tracks compression supervision data, not architecture. All three English-trained compressors show it, deterministic methods show no comparable gap, and the multilingually trained XProvence v1 shows none. Its v2 release, retrained on translated data, empties 92% of Chinese contexts at its aggressive threshold without any warning. Third, in a harder long-context setting, aggressive learned compression drives compressed contexts to or below no-context utility in three of five non-English languages. A translate-then-compress pipeline matches or beats native compression at roughly half the token cost in three of five tested languages. We release all code, compressions, and model outputs. Safe compression budgets are much smaller outside English.
Comments17 pages, 6 figures. Code and artifacts: https://github.com/MantasLukauskas/lost-in-compression