arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Transformer 应该回溯多远?音乐序列模型中的重复与复制

How Far Back Should a Transformer Look? Repetition and Copying in Music Sequence Models

Amir Fathi

arXiv 2610.02837首次发表:更新:

发表机构

University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究符号音乐自回归模型中上下文长度对预测性能的影响,发现长上下文收益主要源于精确复制,并提出了上下文长度测量的方法论检查。

AI 中文摘要

我们研究了在符号音乐的自回归模型中,预测性能如何依赖于模型可用的最大上下文长度,以及长上下文模型利用了哪些信息。我们分别在上下文长度 T 属于 {6, 18, 48, 96, 192, 336} 的情况下训练了一个小型因果 Transformer,并在相同的目标位置上进行评估,评估对象既包括原始词元,也包括非重叠的二元潜码。在 Nottingham 民歌数据集上,词元预测器的测试负对数似然(NLL)在上下文长度从 6 增加到 336 个词元时降低了 71%(每词元 0.986 比特);在长度足够进行相同扫描的 O'Neill 曲调中(302 首测试曲调中的 129 首),该指标也降低了 71%。长上下文带来的大部分收益可以通过精确复制来解释:当目标 16 词元历史的一个较早出现位置进入可用上下文时,改进就会出现;覆盖该出现位置会消除这种收益,而同等规模的不相关破坏则不会;此外,一个简单的复制基线恢复了这个降幅的 94%。一个训练好的 336 词元模型,如果在测试时将其历史限制为最近的词元,同样会失去大部分这种收益。由于许多重复来源于乐谱中展开的书面重复记号,这一结果特别适用于这些渲染后的乐谱表示。相比之下,在 MAESTRO 演奏数据和 MusicNet 乐谱上,精确重复的频率要低得多,降幅也较小(分别为 9% 和 18%),并且大部分收益在 96 个词元时就已经获得。最后,我们审计了先前报告在 16 个词元处出现饱和的草稿,发现了分割泄漏、重叠的潜在感受野、平均预测器以及每文件的节拍网格等问题;我们将这些作为上下文长度测量的方法论检查提出。

英文摘要

We investigate how predictive performance depends on the maximum context available to an autoregressive model of symbolic music, and what information long-context models exploit. A small causal Transformer is trained separately at each context length T in {6, 18, 48, 96, 192, 336} and evaluated on the same target positions, both over raw tokens and over non-overlapping binary latent codes. On Nottingham folk tunes, the token predictor's test NLL decreases by 71% (0.986 bits per token) between 6 and 336 tokens, and by 71% on the O'Neill's tunes long enough for the same sweep (129 of 302 test tunes). Most of the long-context gain is explained by exact copying: the improvement appears when an earlier occurrence of the target's 16-token history enters the available context; overwriting that occurrence removes the gain, whereas equally large unrelated corruption does not; and a simple copy baseline recovers 94% of the reduction. A trained 336-token model likewise loses most of this gain when its history is restricted to recent tokens at test time. Because many repeats arise from written repeat signs expanded in the score, this result pertains specifically to these rendered score representations. By contrast, on MAESTRO performances and MusicNet scores, where exact repeats are substantially less frequent, the reduction is smaller (9% and 18%) and is largely attained by 96 tokens. Finally, an audit of an earlier draft that reported saturation at 16 tokens identified split leakage, overlapping latent receptive fields, an averaging predictor, and a per-file tempo grid; we present these as methodological checks for context-length measurements.

Comments17 pages, 8 figures, 9 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑