实时古尔巴尼跟踪:锡克教基尔坦字幕的基准和参考系统
Live Gurbani Tracking: A Benchmark and Reference System for Captioning Sikh Kirtan
浏览论文内容
中文总结 AI 辅助
该研究针对锡克教基尔坦实时字幕构建基准和参考系统,将任务形式化并组织问题空间,发布含录音和计分器的基准v1,参考系统在最难变体上达57.9%帧准确率,还与基线比较,强调标准指标不适用于此任务,相关成果开源促进改进。
中文摘要 AI 辅助
我们提出了一个用于锡克教基尔坦实时字幕的基准和参考系统,基尔坦是对《斯里古鲁格兰特·萨希卜·吉》经文的连续吟唱。与开放词汇歌词转录不同,基尔坦字幕是一个封闭词汇问题,每行显示必须是经文的准确逐字表述。我们将任务形式化为在每个时间t预测一对(经文ID,行索引)或为空,并将问题空间沿两个正交轴组织成一个2x2矩阵:实时与离线(因果与全音频访问)以及盲与先知(经文身份未知与已知)。我们发布了基准v1版本,包括4个手工标注的基尔坦录音x3个冷启动偏移量 = 12个评估案例,约57分钟的计分音频,以及一个计分器。我们描述了一个参考系统,在最难的变体(实时x盲)上,所有12个案例的总体帧准确率达到57.9%。我们与三个简单基线进行比较,并讨论了为什么标准ASR指标不能衡量此任务所需的显示准确性。基准、参考系统和实时部署均在宽松许可下发布以促进进一步改进。
英文摘要
We present a benchmark and reference system for live captioning of Sikh Kirtan - the continuous, sung recitation of verses from the Sri Guru Granth Sahib Ji (SGGS). Unlike open-vocabulary lyrics transcription, Kirtan captioning is a closed-vocabulary problem: every displayed line must be an exact, word-for-word line from the canonical scripture, because displaying misspelled Gurmukhi is considered religiously inappropriate. We formalize the task as predicting, at every time t, a pair (shabad_id, line_idx) or null, and organize the problem space into a 2x2 matrix along two orthogonal axes: live vs. offline (causal vs. full-audio access) and blind vs. oracle (shabad identity discovered vs. given). We release v1 of the benchmark - 4 hand-annotated Kirtan recordings x 3 cold-start offsets = 12 evaluation cases, ~57 minutes of scored audio - together with a scorer that computes frame accuracy at 1s resolution over a scored region, with a 1s collar and gap-tolerant scoring at segment boundaries. We describe a reference system (fine-tuned 120M IndicConformer -> fuzzy matcher -> state machine; INT8 ONNX; RTF ~0.05 on one Apple Silicon core) that achieves 57.9% overall frame accuracy across all 12 cases (10/12 correct shabad locks) on the hardest variant (live x blind). We compare against three trivial baselines (empty, shifted-5s, perfect) and discuss why standard ASR metrics (WER/CER) measure transcription accuracy rather than the display accuracy this task requires. The benchmark, reference system, and a live deployment are released under permissive licenses to facilitate further improvements.