发表机构
Queen Mary University of London; National Taiwan University; Academia Sinica(伦敦玛丽女王大学; 台湾大学; 中央研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出相对音程网络(RIN)作为音高平滑框架,通过数据驱动的多跳音高差异和$L_1$优化,在语音、歌唱和乐器数据上优于维特比解码,增强鲁棒性。
AI 中文摘要
音高跟踪系统通常将逐帧基频($F_0$)估计器与时间平滑阶段相结合,以获得连续的轨迹。传统的维特比平滑器强制一阶连续性,但缺乏长期时间感知,可能在损坏帧上锁定到八度误差。我们提出了相对音程网络(RIN),一种轨迹平滑框架,它协调逐帧音高估计与数据驱动的多跳音高差异。我们使用可变Q变换互相关提取跨任意帧偏移的鲁棒相对音高间隔。我们将音高平滑表述为$L_1$范数优化问题,并证明其等价于最小费用循环问题,通过线性规划高效求解。在语音、歌唱和乐器数据集上的评估表明,RIN显著改进了弱估计器,在相当的计算成本下匹配或优于维特比解码,并在某些声学退化下提供卓越的鲁棒性。
英文摘要
Pitch tracking systems typically couple a per-frame fundamental frequency ($F_0$) estimator with a temporal smoothing stage to obtain continuous trajectories. Conventional Viterbi smoothers enforce first-order continuity but lack long-term temporal awareness and could lock into octave errors across corrupted frames. We propose Relative Interval Networks (RIN), a trajectory smoothing framework that reconciles per-frame pitch estimates with data-driven multi-hop pitch differences. We extract robust relative pitch intervals across arbitrary frame offsets using Variable-Q Transform cross-correlation. We formulate pitch smoothing as an $L_1$-norm optimization problem and prove its equivalence to a minimum cost circulation problem, solved efficiently via linear programming. Evaluations across speech, singing, and instrumental datasets show that RIN substantially improves weak estimators, matches or outperforms Viterbi decoding at a comparable computational cost, and provides superior robustness under certain acoustic degradation.
CommentsSubmitted to ICASSP 2027