arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06633cs.SD

基于互补音频表示的帧级盘索里调式分类

Frame-Level Pansori Mode Classification with Complementary Audio Representations

Sangheon Park, Seonguk Ju, Suin Chung, Danbinaerin Han, Dasaem Jeong

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对韩国传统声乐盘索里,采用46小时帧级标注与4种互补音频表示,构建调式分类模型,验证其学习调式特征而非记忆曲目,揭示了源分离对线索的影响及预训练的局限性。

中文摘要 AI 辅助

盘索里是韩国传统声乐体裁,其调式体系(조)并非仅由音阶定义,而是音高集合、微分音装饰(시김새)与人声音色的结合。本研究引入了46小时的帧级盘索里调式标注,由专家对全部5种经典바탕进行标注,并评估4种互补输入表示(梅尔频谱图、基频轮廓、MIDI钢琴卷轴及多文化自监督学习编码器),采用两种拆分策略以检测捷径学习。在3种代表性调式中,当保留完整作品时,F1值仅下降2.1至3.6个百分点,表明模型学习到了调式相关特征而非记忆曲目。逐类结果进一步显示,源分离消除了창조依赖的打击乐线索,且通用多文化预训练在Ujo-Gyemyeonjo区分上表现尤其不佳。对跨模态分歧的定性分析揭示了音乐学文献记载的现象,并与已发表的现代创作盘索里的基于乐谱分析结果一致。

英文摘要

Pansori is a traditional Korean vocal genre whose mode system (jo) is defined not by scale alone but by the entanglement of pitch collection, microtonal ornament (sigimsae), and vocal timbre. In this study, we introduce a 46-hour frame-level pansori mode annotation, expert-labeled across all five canonical batang, and evaluate four complementary input representations (mel spectrogram, F0 contour, MIDI piano roll, and a multi-cultural SSL encoder) under two split strategies designed to detect shortcut learning. Across the three well-represented modes, performance degrades by only 2.1--3.6 points of F1 when entire works are held out, indicating that the models learn mode-relevant features rather than memorizing repertoire. Per-class results further show that source separation removes the percussion cue on which changjo depends, and that generic multi-cultural pre-training fails specifically on the Ujo--Gyemyeonjo distinction. Qualitative analysis of cross-modal disagreement recovers musicologically documented phenomena and agrees with published score-based analyses of modern changjak pansori.

补充信息

↑