arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04569cs.CLcs.LG

相关但不完整:指代悬空是硬提示压缩中的范式级失败模式

Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression

Zhengpei Hu, Kai Li, Dapeng Fu, Xuechao Zou, Yuanhao Tang, Yue Li, Tengfei Cao, Jianqiang Huang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究揭示硬提示压缩会导致指代悬空的范式级失败,经实验验证其降低多跳问答准确率,提出自动恢复方法可提升准确率且压缩率仅小幅上升,指出硬压缩器需同时优化相关性与指代完整性。

中文摘要 AI 辅助

硬提示压缩通过对 token、句子或块进行独立评分,并在预算下保留得分最高的单元,以降低长上下文推理成本。我们发现该过程存在一种结构性失败:独立选择会拆分依赖证据对,保留其中一个成员而删除另一个。当保留的文本包含答案但删除的文本定义了解释该答案所需的实体时,我们将此结果称为指代悬空。在压缩率为 0.30 时,Beaver(使用 Qwen3-0.6B 嵌入对连贯块进行排名)在三个多跳问答数据集的桥接示例中,导致 34%-54% 的答案路径不完整。在共享的 HotpotQA 桥接集上,我们测试的所有六种硬压缩器均表现出悬空现象,比例高达 60%;LongBench-v2 单文档 QA 中的每份文档都至少包含一个悬空指代。在使用 Qwen3-8B 评估的悬空示例中,重新插入缺失的支持段落,同时删除非支持段落以保持 token 预算,可将准确率提高 29-34 个百分点(p < 0.0001),至少恢复了保留两个支持段落的上下文所带来的差距的 88%。更强的答案模型无法弥补这种损失:在 MuSiQue 上,GPT-5.5 在压缩上下文上的准确率比保留两个支持段落的上下文低 8.8 个百分点。最后,我们训练了一个紧凑分类器,对省略的句子进行排名,判断它们是否需要解释保留的文本,并在推理时重新插入排名最高的候选,无需支持注释。在 HotpotQA 上使用 Qwen3-8B 时,这种自动恢复将准确率提高了 4.7 个百分点,同时压缩率仅从 0.30 变为 0.31。硬压缩器应同时优化相关性和指代完整性。

英文摘要

Hard prompt compression reduces long-context inference cost by independently scoring tokens, sentences, or chunks and retaining the highest-scoring units under a budget. We identify a structural failure in this procedure: independent selection can split dependent evidence pairs, retaining one member while deleting the other. When retained text contains an answer but deleted text defines the entity needed to interpret it, we call the result referential dangling. At a compression ratio of 0.30, Beaver, which ranks coherent chunks using Qwen3-0.6B embeddings, leaves the answer path incomplete in 34-54% of bridge examples across three multi-hop question answering datasets. On a shared HotpotQA bridge set, all six hard compressors we test exhibit dangling at rates up to 60%, and every document in LongBench-v2 Single-Document QA contains at least one dangling reference. On dangling examples evaluated with Qwen3-8B, reinserting the missing supporting paragraph while removing nonsupporting paragraphs to maintain the token budget improves accuracy by 29-34 percentage points (p < 0.0001), recovering at least 88% of the gap to contexts retaining both supporting paragraphs. Stronger answer models do not absorb the loss: on MuSiQue, GPT-5.5 is 8.8 points less accurate on compressed contexts than on contexts retaining both supporting paragraphs. Finally, we train a compact classifier to rank omitted sentences by whether they are needed to interpret retained text and reinsert the top-ranked candidates without support annotations at inference. On HotpotQA with Qwen3-8B, this automatic restoration improves accuracy by 4.7 points while changing the compression ratio only from 0.30 to 0.31. Hard compressors should optimize both relevance and referential completeness.

补充信息

↑