arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.04021cs.CLcs.LG

当多变为少:语言模型中与位置相关的重复效应

When More Becomes Less: Position-Dependent Repetition Effects in Language Models

Han-yu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究发现语言模型中目标 token 的重复效应与读取位置相关,位移重复会呈现倒U型规律,且该效应经多语言和多模型验证,归因于精确词汇重复而非其他因素。

中文摘要 AI 辅助

完形式探测(Cloze-style probes)会通过改变目标 token 的出现频率来进行研究,这类探测隐含假设是:无论读取槽(readout slot)处于何处,目标的多个副本对预测的影响方式相同。我们证明这一假设不成立。我们的双探测设计将重复目标的前缀固定,仅改变读取位置:相邻探测(adjacent probe)将读取槽置于重复块之后;位移探测(displaced probe)将读取槽置于新句子框架内。相邻重复表现出启动直觉所预测的规律:目标概率P(target)随重复次数N上升并趋于平稳;位移重复则呈现倒U型:P(target)先升至早期峰值,随重复次数增加后下降。在我们测试的13种开源编码器和解码器模型中,位移倒U型表现出逐词下降,其自助法置信区间(bootstrap CI)排除了零值;且在42种多语言单元中,西班牙语、中文、德语和法语的结果均复现了这一现象。一项含6种条件的因果消融实验将该效应隔离为精确的词汇重复,而非长度、通用冗余或语义邻近词暴露。框架语用学控制排除了读取框架的人为因素。内部分析显示,在因果语言模型(causal LMs)中,随N增加,每个目标 token 的注意力下降,而分配给重复块的总预算上升;但在我们探测的掩码语言模型(masked LM)中未出现此情况。改变重复次数的探测不能将读取位置视为与其测量内容正交的变量。

英文摘要

Cloze-style probes that vary how often a target token appears implicitly assume that more copies of a target affect prediction the same way regardless of where the readout slot sits. We show this assumption fails. Our two-probe design holds a repeated-target prefix fixed and varies only the readout position: the adjacent probe places the slot immediately after the repeated block; the displaced probe places it inside a fresh sentence frame. Adjacent repetition behaves as priming intuition predicts: $P(\text{target})$ climbs with $N$ and plateaus. Displaced repetition produces an inverted-U: $P(\text{target})$ rises to an early peak and then declines as more copies are added. The displaced inverted-U shows a per-word drop with bootstrap CI excluding zero in all 13 open-access encoder and decoder models we test, and replicates across Spanish, Chinese, German, and French in 42 of 42 multilingual cells. A six-condition causal ablation isolates the effect to exact lexical repetition rather than length, generic redundancy, or semantic-neighbour exposure. A frame-pragmatics control rules out an artefact of the readout frame. Internally, per-target-token attention falls with $N$ while the total budget assigned to the repeated block grows in causal LMs but not in the masked LM we probe. Probes that vary repetition count cannot treat the readout position as orthogonal to what they measure.

发表机构

  • The University of Hong Kong(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

↑