arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

前沿语言模型在复制方面存在困难:文本可在二维视角下得到更好理解

Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D

Haodong Wen, Yiran Zhang, Yingfa Chen, Kaifeng Lyu

arXiv 2607.16072首次发表:更新:

发表机构

Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现前沿语言模型复制输入字符串存在困难,归因于Transformer架构位置编码问题。为此引入2D-RoPE,将文本组织成二维网格,使复制任务更易学习。实验表明其在复制任务上优势明显,有助于语言建模,鼓励探索二维位置编码潜力。

AI 中文摘要

虽然大语言模型能在数秒内解决高级推理问题,但前沿模型甚至无法完成简单操作:精确复制位于其上下文窗口内的输入字符串。我们将此失败归因于Transformer架构中的位置编码,其归纳偏差倾向于通过基于匹配局部上下文的捷径进行复制,而非仔细定位相应输入位置。为解决此问题,我们引入2D-RoPE,将文本组织成二维网格而非一维序列,并为每个token分配行ID和列ID。在此视角下,复制变为在固定列偏移处检索输入token,使任务易于学习。在合成复制实验中,采用2D-RoPE的浅层Transformer在输入长度比训练时长得数百倍的情况下实现完美复制,而标准位置编码则远远落后。我们进一步表明,2D-RoPE语言模型在复制任务上的优势在高达14亿参数的DCLM大规模预训练中持续存在。总体而言,我们的结果表明二维视角下看待文本有益于语言建模,希望鼓励未来工作进一步探索二维位置编码的潜力。

英文摘要

While large language models (LLMs) can solve advanced reasoning problems in seconds, we show that even frontier models fail to perform a much simpler operation: exactly copying an input string that lies well within their context windows. We attribute this failure to positional encodings in Transformer architectures, whose inductive bias favors copying through a shortcut based on matching local contexts rather than carefully locating the corresponding input positions. To address this issue, we introduce 2D-RoPE, which organizes text into a 2D grid rather than a 1D sequence and assigns each token a row ID and a column ID. Under this view, copying becomes simply retrieving input tokens at a fixed column offset, which makes the task easy to learn. In synthetic copy experiments, shallow Transformers with 2D-RoPE achieve perfect copying at input lengths hundreds of times longer than those seen during training, whereas standard positional encodings fall far behind. We further show that the advantage of 2D-RoPE language models on copy tasks consistently holds in large-scale pretraining on DCLM with model sizes up to 1.4B parameters. Overall, our results suggest that viewing text in 2D can benefit language modeling, and we hope this encourages future work to further explore the potential of 2D positional encodings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑