发表机构
Belmont University(贝尔蒙特大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出分层自监督世界模型,通过Swin V2编码器与条件流匹配模型构建协同音乐创作智能体,提升了和弦与调式检测准确率,可快速生成音乐建议并支持交互演示。
AI 中文摘要
协同音乐智能体需要足够丰富的内部表征以同时支持理解与生成,且需足够灵活以适配人类保留主导权的工作流程。我们提出一种用于符号音乐的分层自监督“世界模型”:一个255万参数的Swin V2编码器,在MIDI钢琴卷帘图像上以JEPA式目标函数(包含音高与时移等变性、掩码嵌入预测及分布正则化)进行训练,未使用任何标签或音乐理论词汇。对冻结嵌入的探测显示,音乐属性可解码的层级与其音乐时间尺度相关:乐句边界可从最粗层级读取,音符密度与和声细节可从最细层级读取。时间与乐句结构仅通过自监督目标函数即可涌现,而和声内容需额外处理;一个小型和弦监督头将联合和弦恢复准确率从0.18提升至0.54,从未受监督的调式检测则从0.16提升至0.70。遵循表征自动编码器范式,条件流匹配模型替代训练好的解码器,在像素空间中从PCA降维的条件输入生成目标窗口,像素F1达0.996;控制变异偏离程度的同层级条件 dropout 还支持掩码修复的图形提示,无需专用修复采样器。该管道在CPU上运行,2.8秒生成建议,在Apple MPS上仅需0.6秒,我们在实时交互演示中展示了这一点。结合基于LLM的“大脑”,这些能力构成了协同音乐创作智能体的核心,服务而非取代人类主导权。
英文摘要
Collaborative music agents need internal representations rich enough to support both understanding and generation, yet flexible enough for a workflow where the human retains agency. We present a hierarchical self-supervised ``world model'' for symbolic music: a 2.55M-parameter Swin V2 encoder trained on MIDI piano-roll images with JEPA-style objectives (pitch- and time-shift equivariance, masked embedding prediction, and a distributional regularizer), using no labels and no music-theory vocabulary. Probing the frozen embeddings shows that the level at which a musical property becomes decodable tracks its musical time scale: phrase boundaries are read off the coarsest levels, note density and harmonic detail off the finest. Temporal and phrase structure emerge from the self-supervised objectives alone, while harmonic content must be asked for; a small chord-supervision head raises joint chord recovery from .18 to .54, and key detection, which is never supervised, from .16 to .70. Following the Representation AutoEncoder paradigm, a conditional flow-matching model stands in for a trained decoder, flowing in pixel space from PCA-reduced conditioning: it reproduces a target window at pixel F1 $0.996$, and the same per-level conditioning dropout that controls how far variations stray also enables graphical prompting for masked inpainting with no inpainting-specific sampler. The pipeline runs on CPU producing a suggestion in $2.8$ s, or $0.6$ s on Apple MPS, which we demonstrate in a live interactive demo. In concert with an LLM-based brain, these capabilities supply the core of a collaborative music creation agent in service of, rather than in place of, human agency.
Comments20 pages, 14 figures. A 6-page version was submitted to the NeurIPS 2026 Creative AI Track. Supplemental website with listening examples: https://drscotthawley.github.io/midi-rae-jepa-son/. Live demo: https://drscotthawley-midi-rae-jepa-son.hf.space/