发表机构
BandLab Technologies(BandLab科技公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于多深度谐波卷积的乐器无关音乐转录模型Harmonica,不同规模变体在准确率或推理效率上表现优异,纳米变体帧F1值优于Basic Pitch,多深度谐波卷积可有效提升转录性能。
AI 中文摘要
本文介绍了Harmonica,这是一类基于多深度谐波卷积构建的乐器无关音乐转录模型。在每个模型规模下,Harmonica在评估的模型中均取得最佳性能:超大模型在乐器无关转录中达到了当前最优性能,而中等规模变体提供了具有竞争力的准确率,且推理速度比所有基线模型更快。为了推动计算效率的极限,纳米变体仅有26.3K参数,运行速度达到实时的1622.5倍,但在开发集上仍达到了0.796的帧F1值,比Basic Pitch高出14.6个百分点。我们还通过与现有谐波聚合方法(包括谐波堆叠、谐波注意力、单深度谐波卷积和HD-Conv层)的对比实验,证明多深度谐波卷积能有效利用谐波信息以提升转录性能。
英文摘要
This paper introduces Harmonica, a family of instrument-agnostic music transcription models built around multi-depth harmonic convolution. At each model scale, Harmonica achieves the best performance among the evaluated models: the x-large model attains state-of-the-art performance in instrument-agnostic transcription, while the medium variant offers competitive accuracy with faster inference than all baselines. Pushing the limit of computational efficiency, the nano variant has only 26.3K parameters and runs at 1,622.5 times real time, yet achieves a frame F1 of 0.796 on the development set, 14.6 percentage points higher than Basic Pitch. We further demonstrate that multi-depth harmonic convolution effectively exploits harmonic information to benefit transcription performance through comparative experiments with existing harmonic aggregation methods, including harmonic stacking, harmonic attention, single-depth harmonic convolution, and the HD-Conv layer.
CommentsSubmitted to ICASSP 2027