arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.28008cs.SD

MIDIBack:通过联合人声-伴奏符号建模实现和声感知的歌唱音高修正

MIDIBack: Harmony-Aware Singing Pitch Correction via Joint Vocal-Accompaniment Symbolic Modeling

Joaquim Cavalcante, Yicheng Gu, Adriel Trajano, Yuri Malheiros, Thais Gaudencio

首次发表
浏览论文内容

中文总结 AI 辅助

提出MIDIBack,一种联合建模人声与伴奏符号序列的音符级歌唱音高修正框架,在多种损坏机制下显著提升原始音高准确率,验证了伴奏上下文的关键作用。

中文摘要 AI 辅助

自动音高修正(APC)需要区分非故意的音准错误与表现性的音高变化。现有系统要么像仅人声方法那样缺乏显式的和声建模,要么不直接使用音符级复调上下文。因此,我们提出MIDIBack,一种音符级APC框架,它在共享的OctupleMIDI序列中联合建模人声和伴奏事件。我们在6种音符损坏机制下评估MIDIBack,包括全局偏移、学习到的音符相关失谐、均匀扰动及其组合。所得模型在整体原始音高准确率(RPA)上达到78.6%,在全局偏移与学习失谐组合下达到81.5%。移除伴奏条件化使RPA在偏移场景下从81.5%降至35.8%,表明伴奏上下文的有效性。关于伴奏移调的案例研究进一步说明人声音符预测

英文摘要

Automatic pitch correction (APC) requires distinguishing the unintended intonation errors from expressive pitch variation. Existing systems either lack explicit harmonic modeling, as vocal-only methods do, or do not directly use the note-level polyphonic context. Therefore, we propose MIDIBack, a note-level APC framework that jointly models the vocal and accompaniment events in a shared OctupleMIDI sequence. We evaluate MIDIBack under 6 note corruption regimes, including global outshift, learned note-dependent detuning, uniform perturbations, and their combinations. The resulting model achieves 78.6% overall raw pitch accuracy (RPA), and 81.5% under combined global outshift and learned detuning. Removing the accompaniment conditioning reduces RPA from 81.5% to 35.8% in outshift, showing the effectiveness of accompaniment context. Case studies on accompaniment modulation further illustrate that vocal note predictions

发表机构

  • Federal University of Paraíba(帕拉伊巴联邦大学)
  • Aalto University(阿尔托大学)
  • Moises AI

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑