arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39929cs.LGcs.CL

RoPE 穷途末路?长上下文失败的理论、诊断与缓解

RoPE at the End of Its Rope? Theory, Diagnosis, and Mitigation of Long-Context Failures

  • University of Illinois at Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)

机构由 AI 辅助整理,请以论文原文为准。

Yuyang Wu, Yufeng Du, Hao Peng

AI总结:

本研究通过允许查询-键尺度不等,完善了 RoPE 长上下文失败理论,提出 RoPE Profiler 诊断工具,并发现推理任务多受语义反转、检索任务多受位置不敏感影响,通过高频重缩放可显著提升准确率。

AI中文摘要:

基于 RoPE 的语言模型的长上下文失败可能源于 RoPE 在维持稳定的词元偏好与区分邻近位置之间的内在权衡。确定要解决哪个弱点以及如何解决,需要对训练后的模型在不同上下文长度下的 RoPE 行为进行更精确的表征。我们通过允许 RoPE 频率上不等的查询-键尺度,解决了先前理论的一个关键局限性,这与实际经验观察非常吻合。我们的理论使两种脆弱性对单个头部和输入均可测量,并量化了高频分量如何支持位置敏感性,同时可能破坏语义稳定性。我们还推导了一个理论上下文长度界限,超过该界限,在指定条件下,固定的注意力分数比较无法同时避免语义反转和位置不敏感。在我们新的理论见解的指导下,我们引入了 RoPE Profiler,一个轻量级、即插即用的诊断工具包,通过重用缓存的查询和键激活,在不增加额外前向传播的情况下增强现有评估。重用评估期间收集的激活,该工具包开销很小。它用两个诊断分数补充标准基准分数,揭示语义和位置弱点,并帮助用户优先处理要解决的方面。至关重要的是,我们在 49 个长上下文任务设置中的评估揭示了一个明显的模式:推理任务主要遭受语义反转,而检索任务主要容易受到位置不敏感的影响。在我们的理论和诊断概况的指导下,有针对性的高频重新缩放无需额外训练即可立即获得收益,在 Qwen3-8B 上将任务准确率提高了最多 20 个百分点,在 Llama-3.1-8B-Instruct 上提高了最多 25 个百分点。

英文摘要:

Long-context failures of RoPE-based language models can arise from RoPE's intrinsic tradeoff between maintaining stable token preferences and distinguishing nearby positions. Determining which weakness to address, and how, requires a more precise characterization of RoPE's behavior in trained models across context lengths. We address a key limitation of prior theory by allowing unequal query-key scales across RoPE frequencies, which aligns well with practical empirical observations. Our theory makes both vulnerabilities measurable for individual heads and inputs, and quantifies how high-frequency components support positional sensitivity while potentially disrupting semantic stability. We also derive a theoretical context-length bound beyond which, under specified conditions, a fixed attention-score comparison cannot jointly avoid semantic reversal and positional insensitivity. Guided by our fresh theoretical insights, we introduce RoPE Profiler, a lightweight, plug-and-play diagnostic toolkit that augments existing evaluations with zero additional forward passes by reusing cached query and key activations. Reusing activations collected during evaluation, the toolkit incurs little overhead. It supplements standard benchmark scores with two diagnostic scores that reveal semantic and positional weaknesses and help users prioritize which aspect to address. Crucially, our evaluations across 49 long-context task settings reveal a distinct pattern where reasoning tasks predominantly suffer from semantic reversal, whereas retrieval tasks are primarily vulnerable to positional insensitivity. Guided by our theory and diagnostic profiles, targeted high-frequency rescaling achieves immediate gains without additional training, improving task accuracy by up to 20 percentage points on Qwen3-8B and 25 percentage points on Llama-3.1-8B-Instruct.

↑