将上下文嵌入整合到富有表现力的MIDI钢琴演奏评估中
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
浏览论文内容
中文总结 AI 辅助
本研究针对MIDI钢琴演奏评估忽略音符依赖、难以聚合表现力属性的问题,结合自监督符号音乐模型Aria和CLaMP3的上下文嵌入,提出适配符号音乐领域的Kernel Audio Distance,发布开源库Pereval。
中文摘要 AI 辅助
对富有表现力的MIDI钢琴演奏的客观评估通常依赖于各个音符的时序、力度、时长等属性统计,但这些方法往往忽略音符之间的依赖关系,这在评估两组演奏的相似性时存在潜在局限。在生成应用中,多样的表现力属性难以聚合为单一标量指标用于模型选择。本研究重新审视了属性范围指标,并探索了来自自监督符号音乐模型Aria和CLaMP3的上下文嵌入的感知属性。我们的听辨研究结果表明,这些模型可作为感知代理,与传统指标相当程度上与单样本人类评分一致。为测量条件分布相似性,我们将Kernel Audio Distance适配到符号音乐领域。与皮尔逊相关和重建误差不同,基于上下文嵌入的核方法无需音符对齐,且对上下文扰动敏感。为便于可复现性,我们发布了开源库Pereval,其整合了包含属性范围指标和深度特征指标的演奏评估工具。
英文摘要
Objective evaluation of expressive MIDI piano performances typically relies on attribute statistics such as timing, velocity, and duration of individual notes. However, these methods often disregard dependencies between notes, which poses a potential limitation in assessing the similarity between two sets of performances. In generative applications, the wide variety of expressive attributes makes it difficult to aggregate them into a single scalar metric for model selection. In this work, we reexamine attribute-scoped metrics and explore the perceptual properties of contextual embeddings from self-supervised symbolic music models, Aria and CLaMP3. Results from our listening study indicate that these models can be used as perceptual proxies, showing agreement with per-sample human ratings on par with traditional metrics. To measure conditional distributional similarity, we adapt Kernel Audio Distance to the symbolic music domain. Unlike Pearson correlation and reconstruction error, kernel-based methods on contextual embeddings do not require note alignment and are sensitive to contextual perturbations. To facilitate reproducibility, we release Pereval, an open-source library that integrates performance evaluation utilities, including both attribute-scoped and deep feature metrics.