arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种语料对齐的奥斯曼体到标准体《古兰经》词映射及确定性诵经校验器

A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator

Yahya Mohamed Elnawasany

arXiv 2609.14967首次发表:更新:

发表机构

Independent Researcher, Egypt(独立研究者,埃及)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对《古兰经》两种正字法差异及诵经校验问题,提出语料对齐词映射与七步规范化流水线,并构建确定性、无LLM的四层匹配诵经校验器,在测试套件上达98.4%准确率。

AI 中文摘要

《古兰经》文本以两种字节级不同的正字法形式流传:一种是每部印刷版穆斯哈夫中使用的奥斯曼体(Uthmani)脚本,另一种是主流阿拉伯语自然语言处理工具所基于的标准(伊姆拉伊,Imla'i)阿拉伯语形式。这一差异集中在一个Unicode字符U+0670(上标阿利夫,superscript alef)上,该字符出现在《古兰经》中一些最常被诵读的词语中,且常被通用阿拉伯语规范化工具静默地错误处理。我们发布了一个包含2,290对、语料对齐的奥斯曼体到标准体词映射,该映射通过对齐完整6,236节经文的两种正字法形式构建而成,并附带一个基于此映射的七步文本规范化流水线。通过该流水线对全部6,236节经文的两种形式进行规范化处理后,90.9%的经文得到完全相同的字符串,我们对其余差异进行了特征化描述,而非断言其已消除。在规范化文本之上,我们构建了一个确定性的、不依赖大语言模型(LLM-free)的《古兰经》诵经校验器,采用四层经文匹配搜索(精确、形态、宽松、模糊)以及跨五个严重级别的词错误率分级反馈。该校验器在发布的测试工具生成的124个案例套件上得分为98.4%(122/124),且两个失败案例共享同一机制:单个替换错误可能使另一节经文成为精确匹配。此外,一项全语料库普查量化了影响16.5%经文的固有纯文本歧义,而在从一套已部署的阿拉伯语自动语音识别(ASR)系统获取的34份诵经转录文本上,该校验器在每种情况下均正确识别出对应经文。我们以开放许可发布该映射、构建脚本、校验器及评估工具;除部署测量(其转录文本不属于我们可发布范围)外,本文中的每个数字均可通过运行它们复现。

英文摘要

Quranic text is distributed in two orthographic forms that are byte-level distinct: the Uthmani script used in every printed mushaf, and the Standard (Imla'i) Arabic form that every mainstream Arabic NLP tool is built for. The gap is concentrated in one Unicode character, U+0670 (superscript alef), which appears in some of the most frequently recited words in the Quran and is silently mishandled by general-purpose Arabic normalizers. We release a 2,290-pair, corpus-aligned Uthmani-to-Standard word mapping constructed by aligning the complete 6,236-verse Quran across both orthographic forms, together with a seven-step text normalization pipeline built on it. Normalizing both forms of all 6,236 verses through that pipeline yields identical strings for 90.9% of verses, and we characterize the residual divergence rather than assert that it is closed. On top of the normalized text, we build a deterministic, LLM-free Quranic recitation validator using a four-layer verse-matching search (exact, morphological, relaxed, fuzzy) and word-error-rate-graded feedback across five severity tiers. The validator scores 98.4% (122/124) on a 124-case suite emitted by the released test harness, and both failures share one mechanism: a single substitution error can make a different verse an exact match. A full-corpus census additionally quantifies an inherent text-only ambiguity affecting 16.5% of verses, and on 34 recitation transcripts drawn from a deployed Arabic ASR system the validator identifies the correct verse in every case. We release the mapping, the script that builds it, the validator, and the evaluation harness under open licenses; every number in this paper except the deployment measurement, whose transcripts are not ours to publish, is reproduced by running them.

Comments6 pages, 4 tables. Dataset, code and evaluation harness: https://github.com/NightPrinceY/muslim-quran-validator

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑