发表机构
Cadi Ayyad University(卡迪·阿雅德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该论文介绍TutlAit数据集,一个众包的摩洛哥塔马齐格特语音语料库,配有阿拉伯语转写和地区口音标签,包含约20.9小时音频,可用于语音识别、翻译和口音识别。
AI 中文摘要
塔马齐格特语(阿马齐格语)与阿拉伯语一起是摩洛哥的两种官方语言之一,然而其在语音技术方面仍然严重资源不足:公开可用的带标签音频稀缺,通常缺乏关于所讲地区变体的信息,且转写质量往往参差不齐。本文介绍了TutlAit数据集,这是一个摩洛哥塔马齐格特语音语料库,配有现代标准阿拉伯语文本和明确的地区口音标签。数据通过TutlAit收集,这是一个专门构建的众包网络应用程序(React 18前端,Django 5 / Django REST Framework后端,PostgreSQL数据库)。通过定向的LinkedIn和Instagram活动招募的母语者创建账户,声明其地区变体(阿特拉斯、苏斯、里夫或其他)和人口统计信息,然后通过两种工作流程做出贡献:文本到音频,即显示一个阿拉伯语句子,志愿者在浏览器中录制其口头塔马齐格特语翻译;以及音频到文本,即播放一段塔马齐格特语摘录,志愿者输入其阿拉伯语转写。一组补充片段从可自由访问的塔马齐格特语视听媒体中获得,使用ELAN进行分段和标注,并通过批量CSV/ZIP管道导入。每次上传都在服务器端转换为16kHz单声道WAV,使用SHA-256进行哈希以拒绝重复,检查时长范围,并由管理员验证。该数据集包含13,384个音频文件,总计75,231秒(约20.9小时,约3.01GB)。阿特拉斯变体占9,956个文件(14.08小时),苏斯变体占3,378个文件(6.75小时);还包括少量里夫(22个文件)和卡拜勒(28个文件)子集。该语料库可用于摩洛哥塔马齐格特语的语音识别、语音翻译和口音识别。
英文摘要
Tamazight (Amazigh) is, together with Arabic, one of the two official languages of Morocco, yet it remains severely under-resourced for speech technology: pub licly available labelled audio is scarce, generally lacks information on the regional variety spoken, and is often of uneven transcription quality. This article describes the TutlAit dataset, a corpus of Moroccan Tamazight speech paired with Modern Standard Arabic text and explicit regional accent labels. The data were collected with TutlAit, a purpose-built crowdsourcing web application (React 18 front end, Django 5 / Django REST Framework back-end, PostgreSQL database). Native speakers recruited through targeted LinkedIn and Instagram campaigns created an account, declared their regional variety (Atlas, Souss, Rif or other) and demographic information, and then contributed through two workflows: Text-to Audio, in which an Arabic sentence is displayed and the volunteer records its oral Tamazight rendering in the browser, and Audio-to-Text, in which a Tamazight excerpt is played and the volunteer types its Arabic transcription. A complemen tary set of segments was obtained from freely accessible Tamazight audiovisual media, segmented and annotated with ELAN and imported through a bulk CSV/ZIP pipeline. Every upload is converted server-side to 16kHz mono WAV, hashed with SHA-256 for duplicate rejection, checked for duration bounds and validated by an administrator. The dataset contains 13,384 audio files totalling 75,231 seconds (approximately 20.9 hours, about 3.01GB). The Atlas variety accounts for 9,956 files (14.08h) and the Souss variety for 3,378 files (6.75h); small Rif (22 files) and Kabyle (28 files) subsets are also included. The corpus can be reused for speech recognition, speech translation and accent identification for Moroccan Tamazight.