arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

塞壬之歌:当近端背景上下文遮蔽远端证据

The Sirens' Song: When Proximal Background Context Overshadows Distant Evidence

Xiaoyu Yang, Jie Lu, Wei Duan, En Yu

arXiv 2609.26718首次发表:更新:

发表机构

Australian Artificial Intelligence Institute (AAII); University of Technology Sydney(澳大利亚人工智能研究院; 悉尼科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文发现长上下文LLM的“近端陷阱”问题,提出LYRA机制重塑注意力分布以增强远端证据检索,并引入ProxBench基准,实验验证了其有效性。

AI 中文摘要

长上下文大语言模型(LLM)专注于从广泛上下文中检索远端证据,然而现有工作主要聚焦于单独克服距离问题。在本工作中,我们识别出“近端陷阱”(Proximity Trap),即对远端证据的关注不足往往并非主要源于距离本身,而是源于与大量任务无关的近端背景的累积竞争。为应对近端陷阱,我们提出LYRA(长上下文重尾相关性对齐,Long-context heavY-tailed Relevance Alignment),一种基于t分布的定向匹配机制,该机制重塑上下文检索分布,将更多注意力权重导向任务相关证据,同时保留编码中的相对位置信息。在LongBench-v2、RULER和LongBench上的大量实验表明,在不同上下文长度和任务类别上均取得了一致的改进。我们进一步引入ProxBench,一个多层级细粒度基准,用于评估在不断增强的近端背景干扰下对远端证据的利用能力。项目页面:此https URL

英文摘要

Long-context LLMs focus on retrieving distant evidence from extensive context, yet existing work has largely focused on overcoming distance alone. In this work, we identify the Proximity Trap, insufficient attention to distant evidence often arises less from distance itself than from cumulative competition with abundant, task-irrelevant proximal background. To address the Proximity Trap, we introduce LYRA (Long-context heavY-tailed Relevance Alignment), a t-distributed directional matching mechanism that reshapes the context retrieval distribution, directing more attention mass toward task-relevant evidence, while preserving the relative positional information encoded. Extensive experiments on LongBench-v2, RULER, and LongBench demonstrate consistent improvements across context lengths and task categories. We further introduce ProxBench, a multi-level fine-grained benchmark for evaluating distant evidence utilization under increasing proximal background interference. Project page: https://xiaoyuyoung.github.io/LYRA/

Comments18 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑