arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于音乐约束序列编辑的无参考歌声音高修正

Reference-Free Singing Pitch Correction via Music-Constrained Sequence Editing

Biao Dong, Jiajun Li, Binzhen Zhu, Mingwei Yi, Tong Liu, Yuanhao Zhang, Jiqing Han, Yongjun He

arXiv 2610.11524首次发表:更新:

发表机构

Faculty of Computing, Harbin Institute of Technology; China Unicom(哈尔滨工业大学计算机学院; 中国联通)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对现有歌声音高修正方法依赖目标旋律或伴奏的问题,提出音乐约束序列编辑方法,在无参考情况下提升了音符级音高准确率,性能优于基线方法。

AI 中文摘要

现有歌声音高修正方法依赖目标旋律或伴奏音轨,而这些在实际场景中可能无法获取。我们将无参考歌声音高修正建模为音乐约束序列编辑任务,仅从输入演唱中确定每个音符是否需要修正以及如何修正。带有轻量歌声领域适配器的预训练符号音乐编码器生成演唱MIDI的上下文表示,基于这些表示,两个轻量修正头通过因子化解构的音高编辑分布联合建模修正必要性与带符号音高修改。从估计的调分布中导出的依赖输入的调性先验随后对候选偏移进行重排序,在无需目标旋律或真实调标注的情况下,偏好符合调性的修正。在真实配对的业余与专业演唱录音上开展的实验显示,该方法将音符级音高准确率从73.25提升至83.43,较基于上下文的基线方法高出3.99个百分点,同时平衡了错误修正与正确演唱音符的保留。

英文摘要

Existing singing pitch correction approaches rely on target melodies or accompaniment tracks, which may be unavailable in practice. We formulate reference-free singing pitch correction as a music-constrained sequence editing task that determines whether and how each note should be corrected from the input performance alone. A pretrained symbolic music encoder with lightweight singing-domain adapters produces contextual representations of the singing MIDI. Based on these representations, two lightweight correction heads jointly model correction necessity and signed pitch modification through a factorized pitch-editing distribution. An input-dependent tonal prior derived from the estimated key distribution then reranks the candidate offsets, favoring tonally compatible corrections without target melodies or ground-truth key annotations. Experiments on real paired amateur and professional singing recordings show that the method improves note-level pitch accuracy from 73.25 to 83.43, outperforming a context-based baseline by 3.99 percentage points while balancing error correction and preservation of correctly performed notes.

CommentsSubmitted to ICASSP 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑