基于VQ-VAE的音高轮廓分词及其在韩国传统音乐分析中的应用
Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis
AI总结:
该研究针对连续音高运动的音乐传统,用VQ-VAE从无标注音频学习音高轮廓分词,在韩国传统音乐分析中验证了分词可作为语料库级分析单元的有效性。
AI中文摘要:
音乐的计算分析常依赖离散表示,但许多音乐传统以连续的音高运动为核心,难以分割为类似音符的单元。对于这类传统,分析所需的离散单元并非预先给定。我们通过从无标注音频中直接学习局部音高轮廓模式的词汇表来解决这一问题,使用VQ-VAE将固定长度的轮廓片段量化为有限码本。为使学习到的分词在分割位置及时序、音高范围的微小变化中保持稳定,我们采用在一组候选时域和频域变换的最佳对齐下评估的重建目标训练模型。将其应用于韩国传统音乐时,学习到的分词在无监督情况下恢复了专家定义的sigimsae类别的信息;在盘索里(pansori)中,单个分词与两种主要调式——鸡鸣调(Gyemyeonjo)和羽调(Ujo)对齐,支持将其作为以轮廓为核心的传统音乐语料库级分析的单元。
英文摘要:
Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.