arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.10233cs.SDcs.LG

MeloBottleneck:基于潜在子序列瓶颈的自监督旋律骨架提取

MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck

Fan Bu, Rongfeng Li, Linfeng Fan

AI总结:

研究旋律骨架提取问题,提出MeloBottleneck自监督框架,通过特定算子和训练方式生成旋律骨架,在多种模式下评估,结果显示该方法比伪标签模仿迁移性更强,还能提升基于BM25的片段检索效果。

AI中文摘要:

旋律骨架提取旨在去除装饰音,得到保留结构音符的较短旋律。现有方法依赖手工制作的简化规则或用启发式或程序生成的伪标签训练的音符显著性分类器,存在偏差且未明确优化连贯的简化旋律。本文介绍了MeloBottleneck自监督框架,通过硬瓶颈提取器选择音符事件,节奏闭合算子生成自洽骨架,重新装饰解码器重建输入旋律。训练结合了重建、冻结的自回归旋律先验、程序装饰视图间的装饰不变一致性和装饰排除。在三种模式下评估,结果表明学习骨架作为潜在子序列比伪标签模仿产生更强大的迁移。

英文摘要:

Melody skeleton extraction aims to derive a shorter melody that preserves structural notes while removing ornaments. Prior methods rely on hand-crafted reduction rules or note-wise salience classifiers trained with heuristically or procedurally generated pseudo-labels. Such supervision can inherit generator bias and does not explicitly optimize a coherent reduced melody. We introduce MeloBottleneck, a self-supervised framework that represents a skeleton as a length-controlled, order-preserving latent subsequence. A hard-bottleneck extractor selects note events, a rhythmic-closure operator produces a self-consistent skeleton, and a re-ornamentation decoder reconstructs the input melody. Training combines reconstruction, a frozen autoregressive melody prior, ornament-invariant consistency across procedurally ornamented views, and ornament exclusion. We evaluate three regimes: synthetic out-of-distribution ornament-to-skeleton, TAVERN variation-to-theme, and Jiugong ornamented-to-gongche. A matched pseudo-label classifier excels on the synthetic benchmark, while MeloBottleneck transfers better, achieving competitive selection quality on TAVERN and Jiugong. Skeletonized melodies also improve BM25-based fragment retrieval, boosting Recall@K and MRR while reducing query time. Overall, the results suggest that learning skeletons as latent subsequences yields more robust transfer than pseudo-label imitation.

补充信息

↑