arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Music-JEPA:从动作中学习声音的世界模型

Music-JEPA: Learning a World Model of Sound from Action

Ziyu Wang, Kun Fang, Yann LeCun

arXiv 2607.22000首次发表:更新:

AI 中文总结

研究提出Music-JEPA,将音乐构建为动作条件系统,用JEPA学习钢琴声音世界模型,在离线状态下用配对数据训练。实验表明模型能捕捉音乐动作与声音关系,其表示支持下游任务并能通过规划实现钢琴转录。

AI 中文摘要

联合嵌入预测架构(JEPA)最近作为一种通过预测潜在表示来学习世界模型的范式出现,为自监督学习提供了一个有前景的方向。虽然最初已尝试将JEPA应用于音乐领域,但尚不清楚此类框架如何自然地支持音乐世界模型的形成。在这项工作中,我们建议通过将音乐构建为动作条件系统,使用JEPA学习钢琴声音的世界模型:音频被视为状态,钢琴卷帘被视为乐器动作。给定当前音频状态和动作,模型预测未来的音频状态,反映人类通过互动学习音乐声音的方式。该模型在完全离线设置下使用配对的音频 - 钢琴卷帘数据进行训练。实验表明,学习到的模型捕捉到了音乐动作与其产生的声音之间的关系。所得表示支持下游任务,包括节拍跟踪、作曲家识别和调估计,并通过搜索最能解释目标声音的动作,通过规划实现钢琴转录。

英文摘要

Joint Embedding Predictive Architectures (JEPA) have recently emerged as a paradigm for learning world models by predicting latent representations, offering a promising direction for self-supervised learning. While initial attempts have applied JEPA to the music domain, it remains unclear how such frameworks can naturally support the formation of a world model for music. In this work, we propose to learn a world model of piano sound using JEPA by framing music as an action-conditioned system: the audio is treated as the state, and the pianoroll as the instrument action. Given a current audio state and an action, the model predicts the resulting future audio state, mirroring how humans learn musical sound through interaction. The model is trained in a fully offline setting using paired audio-pianoroll data, without environment interaction. Experiments show that the learned model captures the relationships between musical actions and their resulting sound. The resulting representations support downstream tasks, including beat tracking, composer identification, and key estimation, and enable piano transcription via planning, by searching for actions that best explain a target sound.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑