arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03419cs.SDcs.AIcs.CVcs.MMeess.IV

多任务多帧视觉钢琴转录

Multi-Task Multi-Frame Visual Piano Transcription

Yonghyun Kim, Hoyeol Sohn, Juhan Nam, Alexander Lerch

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉钢琴转录系统音尾准确率低、未报告力度的问题,提出V2N多任务视觉钢琴转录系统,采用逐帧监督训练,在PianoVAM和R3上取得最优结果。

中文摘要 AI 辅助

基于音频的钢琴转录在音头、音高和力度方面表现良好,但延音踏板会让声音在琴键松开后持续存在,因此音频系统预测的是踏板延伸的音尾而非物理琴键的松开。然而现有的视觉钢琴转录(VPT)系统仅关注短视频窗口内的音头检测,音尾准确率远落后于音头,且未报告音符级力度。为解决这些不足,我们提出首个完整的VPT系统V2N(Video to Notes):共享时间骨干网络为音头、音尾、琴键保持和力度提供任务特定的输出头,采用逐帧监督而非仅窗口中心的方式联合训练。 ablation实验表明多任务监督可实现音尾和力度预测,同时提升音头准确率;更长的时间上下文能进一步改善性能。V2N在PianoVAM和R3数据集上取得了新的最优结果。

英文摘要

Audio-based piano transcription performs well on onset, pitch, and velocity, but the sustain pedal lets sound persist long after key release, so audio systems predict pedal-extended offsets rather than physical key release. Yet existing Visual Piano Transcription (VPT) systems focus on onset detection from short video windows, offset accuracy lags onset by a wide margin, and note-level velocity has not been reported. To address these gaps, we present V2N (Video to Notes), the first complete VPT system: a shared temporal backbone feeds task-specific heads for onset, offset, key hold, and velocity, jointly trained with per-frame supervision rather than only at the window center. Ablations show that multi-task supervision enables offset and velocity prediction while improving onset accuracy; longer temporal context yields further improvements. V2N sets new state-of-the-art results on PianoVAM and R3.

补充信息

↑