arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越钢琴:基于可Differentiable SoundFont代理的跨乐器MIDI力度估计

Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies

Zhanhong He, Hanyu Meng, David Defeng Huang, Roberto Togneri

arXiv 2608.08985首次发表:更新:

AI 中文总结

本文针对标签稀缺场景下的跨乐器MIDI力度估计问题,提出基于可微分SoundFont代理(Diff-SFProxy)的方法,在钢琴和吉他实验中验证其有效性,优于波形域可微分合成器(Diff-Synth)。

AI 中文摘要

许多音乐数据集包含MIDI音符但缺乏可靠的力度值,默认采用固定值。这种缺失在钢琴以外的乐器领域尤为突出,因为力度是表现力渲染、音乐生成和演奏分析的核心组成部分。本文研究标签稀缺场景下的跨乐器MIDI力度估计问题。从经钢琴训练的力度估计器出发,我们将目标乐器适配重构为预测渲染器条件下的力度,其渲染结果需匹配演奏音频的动态特性。该适配可由可微分合成器(Diff-Synth)或我们提出的可微分SoundFont代理(Diff-SFProxy)驱动。我们重点介绍Diff-SFProxy:它通过逐音符、与响度相关的声学参数而非波形重建来监督力度,将梯度聚焦于力度依赖的行为。在钢琴和吉他上的实验表明,Diff-SFProxy对跨乐器MIDI力度估计有效,而波形域的Diff-Synth会降低性能。

英文摘要

Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper studies cross-instrument MIDI velocity estimation in this label-scarce setting. Starting from a piano-trained velocity estimator, we recast target-instrument adaptation as predicting renderer-conditioned velocities whose rendering matches the dynamics of the performance audio. This adaptation can be driven by either differentiable synthesizers (Diff-Synth) or our proposed differentiable SoundFont proxies (Diff-SFProxy). We highlight the Diff-SFProxy: it supervises velocity through note-wise, loudness-related acoustic parameters rather than waveform reconstruction, focusing gradients on velocity-dependent behavior. Experiments on piano and guitar show that Diff-SFProxy is effective for cross-instrument MIDI velocity estimation, while waveform-domain Diff-Synth degrades performance.

CommentsAccepted to ISMIR2026 conference

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑