AI 中文总结
TART提出模块化四阶段音频到指谱流水线,解决吉他转录中表现性技术捕捉、弦品分配和噪声泛化问题,在四个基准上显著优于基线。
AI 中文摘要
吉他自动音乐转录(AMT)仍受三个挑战的限制:现有系统往往无法捕捉滑音、推弦和打击性击弦等表现性技术;它们经常将音符分配到错误的弦-品组合;并且它们通常在干净录音上训练,限制了其对嘈杂真实世界音频的泛化能力。为解决这些挑战,我们提出了TART,一个模块化的四阶段音频到指谱流水线,包括(1)音频到MIDI转录模型,(2)表现性技术分类器,(3)用于弦-品分配的音频条件T5编码器-解码器,以及(4)自动指谱生成器。我们在GuitarSet、EGDB以及两个增强基准Noisy GuitarSet和Noisy EGDB上以零样本设置评估TART。在这四个基准上平均,TART实现了81.35%的音频到MIDI F50(比最佳先前基线高6.67个百分点),71.8%的弦-品Tab F1(比最佳先前基线高8.5个百分点),以及54.08%的端到端Tab F1。据我们所知,TART是第一个直接从吉他音频生成带有指法和表现性技术注释的吉他指谱的框架。
英文摘要
Automatic Music Transcription (AMT) for guitar remains limited by three challenges: existing systems often fail to capture expressive techniques such as slides, bends, and percussive hits; they often assign notes to incorrect string-fret combinations; and they are typically trained on clean recordings, limiting their generalization to noisy real-world audio. To address these challenges, we propose TART, a modular four-stage audio-to-tablature pipeline consisting of (1) an audio-to-MIDI transcription model, (2) an expressive technique classifier, (3) an audio-conditioned T5 encoder-decoder for string-fret assignment, and (4) an automated tablature generator. We evaluate TART in a zero-shot setting on GuitarSet, EGDB, and two augmented benchmarks, Noisy GuitarSet and Noisy EGDB. Averaged across these four benchmarks, TART achieves 81.35% audio-to-MIDI F50 (+6.67 points over the best prior baseline), 71.8% string-fret Tab F1 (+8.5 points over the best prior baseline), and 54.08% end-to-end Tab F1. To our knowledge, TART is the first framework to generate guitar tablature with both fingering and expressive technique annotations directly from guitar audio.
CommentsISMIR 2026