发表机构
Leiden University; Leiden Institute of Advanced Computer Science (LIACS)(莱顿大学; 莱顿高级计算机科学研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现场伴奏中节奏漂移问题,提出静默节拍器(SiMe),通过周期性编码节拍相位作为独立条件通道提供时间参考,使节拍对齐提升3.2倍,并超越非因果基线。
AI 中文摘要
现场伴奏模型为传入的音频流生成音乐,在听到后续内容之前便对每个输出帧做出承诺。在这种严格因果的设置中,模型必须从自身不完美的过去中推断出速度、节拍和节拍相位,由此累积的误差会迅速以节奏漂移的形式变得可闻。简而言之,模型有耳朵但没有时间参考,因此当耳朵听到不完美、模糊的音乐时,模型会产生有缺陷的输出。我们提出了静默节拍器(SiMe),它提供了时间参考,将节拍内和小节内的相位编码为周期函数,将其与速度和拍号配对,并将结果作为单独的条件通道提供。由于该参考独立于生成的音频,它不会漂移。互补的辅助头塑造了潜在表示,其中包括一个预测模型自身未来令牌的新颖辅助头。当节拍信号取自真实标注时,节拍对齐比严格因果基线提高了3.2倍,并超过了拥有一整秒前瞻的非因果参考。输入与伴奏之间的连贯性保持在参考值的单点以内。这些结果表明,流式伴奏系统应像人类合奏那样将节奏视为要共享的信号,而非推断的信号。
英文摘要
Live accompaniment models generate music for an incoming audio stream, committing to each output frame before hearing what comes next. In this strictly causal setting the model must infer tempo, meter, and metrical phase from its own imperfect past, whereby compounding errors quickly become audible as rhythmic drift. Put simply, the model has ears but no temporal reference, so when the ears hear imperfect, ambiguous music, the model will produce a flawed output. We propose Silent Metronome (SiMe), which gives it the temporal reference, encoding the phase within the beat and within the bar as periodic functions, pairing them with tempo and time signature, and supplying the result as a separate conditioning channel. Because this reference is independent of the generated audio, it cannot drift. Complementary auxiliary heads shape the latent representation, including a novel head that predicts the model's own future tokens. With the metrical signal taken from ground-truth annotations, beat alignment improves by a factor of 3.2 over the strictly causal baseline and surpasses a non-causal reference granted a full second of look-ahead. Coherence between input and accompaniment stays within a single point of that reference. These results suggest that streaming accompaniment systems should treat rhythm as a signal to be shared, as human ensembles do, rather than inferred.
Comments5 pages, 2 figures, 1 table