arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30314eess.AS

基于自适应动态节奏语言模型(ADRM)的弱监督塔布拉鼓击奏转录

Weakly Supervised Tabla Stroke Transcription via an Adaptive Dynamic Rhythm Language Model (ADRM)

Rahul Bapusaheb Kodag, Vipul Arora

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对弱监督塔布拉鼓击奏转录任务,提出结合CTC声学模型与ADRM的框架,发布相关数据集,实验证实ADRM可显著降低击奏错误率。

中文摘要 AI 辅助

塔布拉鼓击奏转录(TST)是印度斯坦音乐节奏结构分析的核心环节,但由于其复杂且动态的节奏组织形式以及强标注数据的稀缺性,该任务仍具挑战性。现有方法大多依赖带有音素级标注的全监督学习,此类标注成本高昂且难以规模化应用。本研究针对弱监督场景下的TST问题,仅使用无音素时间对齐的符号击奏序列开展研究。我们提出一种框架,将基于连接时序分类(CTC)的声学模型与序列级节奏语言模型相结合以进行重评分,该架构与自动语音识别中的类似设置相近。声学模型生成解码格,随后通过自适应动态节奏语言模型(ADRM)进行优化,ADRM将以$t\bar{a}la$为条件的符号节奏规律与局部击奏动态相结合。此外,我们发布了一个新的演奏记录塔布拉鼓数据集,命名为Tabla Improvisation Dataset,以及一个互补的合成数据集,用于序列级弱监督TST。实验表明,与仅使用声学模型解码相比,采用ADRM可显著且持续降低击奏错误率,这证实了在解码格重评分过程中融入符号节奏规律对实现准确转录的益处。

英文摘要

Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani music, yet it remains challenging due to complex and dynamic rhythmic organization and the scarcity of strongly annotated data. Existing approaches largely rely on fully supervised learning with onset-level annotations, which are costly and impractical at scale. This work addresses TST in a weakly supervised setting, using only symbolic stroke sequences without temporal alignment of onsets. We propose a framework that combines a Connectionist Temporal Classification (CTC)-based acoustic model with a sequence-level rhythmic language model for rescoring, similar to that used in automatic speech recognition. The acoustic model produces a decoding lattice, which is refined using an Adaptive Dynamic Rhythm Language Model (ADRM) that combines $t\bar{a}la$-conditioned symbolic rhythmic regularities with local stroke dynamics. Moreover, we release a new performance-recorded tabla dataset, named \emph{Tabla Improvisation Dataset}, along with a complementary synthetic dataset for sequence-level weakly supervised TST. Experiments demonstrate consistent and substantial reductions in stroke error rates with ADRM compared to those with acoustic-only decoding, confirming the benefit of incorporating symbolic rhythmic regularities during lattice rescoring for accurate transcription.

↑