AI 中文总结
本文针对标签稀缺场景下的跨乐器MIDI力度估计问题,提出基于可微分SoundFont代理(Diff-SFProxy)的方法,在钢琴和吉他实验中验证其有效性,优于波形域可微分合成器(Diff-Synth)。
AI 中文摘要
许多音乐数据集包含MIDI音符但缺乏可靠的力度值,默认采用固定值。这种缺失在钢琴以外的乐器领域尤为突出,因为力度是表现力渲染、音乐生成和演奏分析的核心组成部分。本文研究标签稀缺场景下的跨乐器MIDI力度估计问题。从经钢琴训练的力度估计器出发,我们将目标乐器适配重构为预测渲染器条件下的力度,其渲染结果需匹配演奏音频的动态特性。该适配可由可微分合成器(Diff-Synth)或我们提出的可微分SoundFont代理(Diff-SFProxy)驱动。我们重点介绍Diff-SFProxy:它通过逐音符、与响度相关的声学参数而非波形重建来监督力度,将梯度聚焦于力度依赖的行为。在钢琴和吉他上的实验表明,Diff-SFProxy对跨乐器MIDI力度估计有效,而波形域的Diff-Synth会降低性能。
英文摘要
Many music datasets contain MIDI notes but lack reliable velocities, defaulting to a constant value. This absence is especially problematic outside the piano domain, as velocity is a core component for expressive rendering, music generation, and performance analysis. This paper studies cross-instrument MIDI velocity estimation in this label-scarce setting. Starting from a piano-trained velocity estimator, we recast target-instrument adaptation as predicting renderer-conditioned velocities whose rendering matches the dynamics of the performance audio. This adaptation can be driven by either differentiable synthesizers (Diff-Synth) or our proposed differentiable SoundFont proxies (Diff-SFProxy). We highlight the Diff-SFProxy: it supervises velocity through note-wise, loudness-related acoustic parameters rather than waveform reconstruction, focusing gradients on velocity-dependent behavior. Experiments on piano and guitar show that Diff-SFProxy is effective for cross-instrument MIDI velocity estimation, while waveform-domain Diff-Synth degrades performance.
CommentsAccepted to ISMIR2026 conference