什么算错误?在《古兰经》背诵转录中标注背诵事件
What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts
- LemoniLab FZCO
- Innopolis University(因诺波利斯大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究标注《古兰经》背诵转录中的错误事件,提出可执行评估器,发现基线失败源于标注界面问题,并指出剩余挑战是事件边界约定而非检测。
AI中文摘要:
从ASR转录文本检查《古兰经》背诵,需要区分未解决的错误与重复、修正、开头用语和被接受的拼写差异。我们报告了对100个生产录音案例的完整人工标注:348个评分单元和162个局部事件,涵盖十个组合标签。一个可执行的评估器同时评分标签和词位置。简单的diff达到标签感知F1 0.525和定位F1 0.826;适配的生产清洗/对齐组件达到0.518和0.786,两者的精确跨度F1均为0.505。修正适配器的词坐标后,恢复了全部五个标注的重复事件,这表明在解释基线失败之前必须检查标注界面。在初步试点中,三个编码代理和八个模型进行的八次单次20分钟运行,标签感知F1范围从0.143到0.892:七次远高于所有基线,一次因缺少归一化步骤而低于朴素diff。在这六次运行中,972个黄金事件实例中有970个获得了重叠预测,因此剩下的不是检测问题而是约定问题:跨度范围,以及边界由裁决规定而非文本可见的标签。162个事件中有7个击败了所有六次同日运行,其中五个是一个拼写规则,最强的运行仍然漏掉了这些事件。没有一次运行在构建之前进行标注,因此试点仅衡量了任务中算法的一半。
英文摘要:
Checking Quran recitation from an ASR transcript requires distinguishing unresolved mistakes from repetitions, repairs, opening formulas and accepted spelling differences. We report a completed human annotation of 100 production recording cases: 348 scored units and 162 localized events across ten combined labels. An executable evaluator scores labels and word positions together. A plain diff reaches label-aware F1 0.525 and localization F1 0.826; adapted production cleaner/alignment components reach 0.518 and 0.786, with exact-span F1 0.505 for both. Correcting the adapter's word coordinates recovers all five annotated repetition events, showing why annotation interfaces must be checked before interpreting baseline failures. In a preliminary pilot, eight single 20-minute runs across three coding agents and eight models span label-aware F1 0.143 to 0.892: seven land far above every baseline, and one collapses below the naive diff from a missing normalization step. Across the six, 970 of 972 gold-event instances draw an overlapping prediction, so what remains is not detection but convention: span extent, and the labels whose boundary is stipulated by adjudication rather than visible in the text. Seven of 162 events defeat all six same-day runs, five of them one orthographic rule, and the strongest run still misses the same ones. No run annotated before building, so the pilot measures the algorithm half of the task only.