Neyshekar:用于自动语音识别的开放波斯语朗读语音语料库
Neyshekar: An Open Persian Read-Speech Corpus for Automatic Speech Recognition
浏览论文内容
中文总结 AI 辅助
Neyshekar 是一个开放波斯语朗读语音语料库,覆盖正式与非正式语言,含 99 小时验证录音,相比 Common Voice 显著降低 WER,并公开代码与数据。
中文摘要 AI 辅助
Neyshekar 作为一个开放的波斯语朗读语音语料库被提出,旨在覆盖正式与非正式语言、命名实体以及更长的语句。在第 6 版中,提供了来自 190 位贡献者的 62,279 条经过验证的录音,总计 99.02 小时,包含 34,541 条不同的已记录提示。提示池由人工撰写的材料、语境化的同形异义词以及经过审查的语言模型生成的文本组成。文本条目使用 shekar 库进行规范化,该库支持正式和非正式的波斯语,并且每份提交的录音都根据统一的验证标准进行了审查。约 24% 的已发布片段被自动分类器归类为非正式;这些语域标签未经人工验证。提供了条目级评分者标签以实现可复现的一致性估计,不透明的每片段贡献者标识符使说话者不重叠的划分可审计,并支持贡献者聚类的置信度估计,此外还包含一个文本不重叠的测试子集,用于评估超出先前所见提示的性能。每个已发布片段的每位贡献者录音负载和免参考信号质量均被表征。在共享处理流程下,语料库特性与波斯语 Common Voice 进行了比较。通过两种 ASR 架构、三种优化种子、WER 和 CER,以及在公共 PSRB 样本上的独立评估来评估实用性。与约 32 小时的时长匹配的 Common Voice 训练相比,Whisper 的域内 WER 降低了 9.5 个百分点,XLS-R 降低了 11.6 个百分点,两种架构在独立 PSRB 样本上均降低了约八个百分点。迁移和混合带来的收益在不同架构和训练预算下并不一致。该语料库以 CC0 协议发布;代码和数据通过项目仓库在此 https URL 提供。
英文摘要
Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utterances. In version 6, 62,279 validated recordings totalling 99.02 hours are provided from 190 contributors, with 34,541 distinct recorded prompts. The prompt pool was assembled from human-written material, contextualised homographs, and reviewed language-model-generated text. Text entries were normalised with the shekar library, which supports both formal and informal Persian, and every submitted recording was reviewed against a common validation rubric. About 24% of released clips are classified as informal by an automatic classifier; these register labels are not human-validated. Item-level rater labels are provided for reproducible agreement estimation, opaque per-clip contributor identifiers make the speaker-disjoint partitioning auditable and support contributor-clustered uncertainty estimates, and a text-disjoint test subset is included for evaluation beyond previously seen prompts. Per-contributor recording load and reference-free signal quality are characterised for every released clip. Corpus characteristics are compared with Persian Common Voice under shared processing. Utility is assessed through two ASR architectures, three optimisation seeds, WER and CER, and independent evaluation on the public PSRB sample. Against duration-matched Common Voice training at approximately 32 hours, in-domain WER is reduced by 9.5 points for Whisper and 11.6 points for XLS-R, and by approximately eight points for both architectures on the independent PSRB sample. Transfer and mixture benefits are not consistently observed across architectures and training budgets. The corpus is released under CC0; code and data are made available through the project repository at https://github.com/amirivojdan/neyshekar.
发表机构
- Shekar AI
机构由 AI 辅助整理,请以论文原文为准。