Agogic:适用于大语言模型原生文本到符号音乐生成的性能计时音乐 token
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
浏览论文内容
中文总结 AI 辅助
该研究固定模型规模、数据等变量,发现音乐 token 化的表示形式而非模型规模是影响文本到符号音乐生成分布保真度的关键,发布了相关工具链、模型及语料库,为该领域提供了可测量的表示形式研究基础。
中文摘要 AI 辅助
文本到音乐语言模型通常默认面临一个选择:如何对音乐进行 token 化。该过程通常与主干模型、数据和训练方案紧密关联,其影响从未被单独测量过。我们固定预训练的 Qwen3.5(参数规模为 0.8B 至 27B)、数据、计算预算和解码方式,仅在七种 token 化方案之间交换表示形式,并将纹理指标锚定到每种表示形式的无模型上限。结果的排序清晰且令人惊讶:表示形式而非模型规模是分布保真度的关键变量。将主干模型缩放 34 倍几乎不会改变 Frechet 音乐距离(FMD),而切换表示形式可使 FMD 减半。我们发布的 PMT(性能分辨率流,具备 10 毫秒计时、每个音符力度、多轨纹理,共 609 个符号)在 0.8B 参数规模下达到 FMD 159,而 beat grids 的 FMD 为 272-286(低 1.7-1.8 倍,其他地方最高达 2.8 倍;非重叠自助置信区间),因此 0.8B 参数规模的性能分辨率模型优于 27B 参数规模的 beat grid。该结果在从头训练的 26M 参数主干模型和第二个性能分辨率 tokenizer 上也成立:这是该类方法的属性,而非幸运词汇表的结果。这也不是更细网格的人工产物:将 PMT 的起始时间对齐到 beat grids 的分辨率后,其 FMD 仍比两者高 67-129(样本量 n=500)。该效果属于分布层面;其是否可被感知是另一个问题,我们的探究未作解答,相关人类研究已预注册。原生字幕贴合度较弱但可分离:轻量级解码时约束使乐器 F1 翻倍(从 0.28 升至 0.60)、正确调式比例翻倍(从 0.16 升至 0.35),且无分布层面损失。我们发布了工具链、25 个以上检查点、两个语料库(86.6k 个跨字幕/MIDI/ABC/音频的对齐样本;625 万个带字幕样本,是目前最大的音乐相关语料库),以及一种 imprinting 诊断方法:已发布的文本到 MIDI 系统在不相交领域中,其训练分布对字幕的响应近乎不变(和弦时间占比 72% 对 71%)。现在,该领域的下一个表示形式主张可被测量,而非仅被断言。
英文摘要
Text-to-music language models begin with a choice usually made by default: how to tokenize music. Normally entangled with backbone, data, and recipe, its effect has never been measured in isolation. We fix pretrained Qwen3.5 (0.8B-27B), data, budget, and decoding, and swap only the representation across seven tokenizations, anchoring texture metrics to each representation's model-free ceiling. The ordering is clean and surprising: representation, not model size, is the binding variable for distributional fidelity. Scaling the backbone 34x barely moves Frechet Music Distance (FMD), whereas switching representation halves it. PMT, a performance-resolution stream we release (10 ms timing, per-note velocity, multi-track texture; 609 symbols), reaches FMD 159 at 0.8B against 272-286 for beat grids (1.7-1.8x lower, up to 2.8x elsewhere; non-overlapping bootstrap CIs), so a 0.8B performance-resolution model beats a 27B beat grid. It reappears on a 26M from-scratch backbone and a second performance-resolution tokenizer: a property of the class, not one lucky vocabulary. Nor is it a finer-lattice artifact: snapping PMT's onsets to the beat grids' resolution still leaves it 67-129 FMD ahead of both (n=500). The effect is distributional; whether it is audible is a separate question, left open by our probe, with a human study pre-registered. Native caption adherence is weak but separable: a lightweight decode-time constraint doubles instrument-F1 (.28 to .60) and Correct-Key (.16 to .35) at no distributional cost. We release the harness, 25+ checkpoints, two corpora (86.6k aligned across caption/MIDI/ABC/audio; 6.25M captioned, the largest for music), and an imprinting diagnostic: published text-to-MIDI systems reproduce their training distribution near-invariant to the caption (72% vs. 71% chord-time on disjoint domains). The field's next representation claim can now be measured, not asserted.