arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.25621cs.SDcs.AI

不协和频谱(Dissonance Spectrum,DS):显式建模感知频率交互以实现更好的音乐理解

Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding

  • Beijing Institute for General Artificial Intelligence(北京通用人工智能研究院)
  • Central Conservatory of Music(中央音乐学院)
  • Peking University(北京大学)

机构由 AI 辅助整理,请以论文原文为准。

Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen, Sirui Zhang, Haoxin Zhang, Xin Jin, Duo Xu, Xiaobing Li, Song-Chun Zhu

AI总结:

本文提出不协和频谱(DS)这一时频表示,经实验验证其在音乐相关任务中性能优于基线等方法,可作为音乐理解的可解释互补表示。

AI中文摘要:

传统音乐表示描述了能量随时间和频率的变化,但未显式呈现同时频率分量间的关系。我们提出不协和频谱(Dissonance Spectrum,DS),这是一种非负时频表示,它将基于容忍度的有理音高关系核(带有对数谐波距离)应用于恒定Q谱,并将聚合的成对交互归因于各个频率仓。受控音乐理论测试显示,其在音程、和声功能关联及教会调式上具有强序数一致性,在不同和弦转位上则有较弱但显著的一致性。随后,DS由一个轻量并行分支编码,该分支的零初始化残差投影在初始化时保留基线函数。在开放域音乐问答、分类及维度音乐情感识别的6组配对训练种子中,DS在所有报告的端点上均取得了比未改变基线、参数匹配的高斯输入分支及架构匹配的幅度-CQT分支更高的均值。这些结果支持DS作为一种可解释的互补表示,而听众特定感知及更广泛任务覆盖仍是未解决的问题。

英文摘要:

Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time--frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and weaker but significant agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems.

↑