SHAMS:黎凡特阿拉伯语音频发音基准
SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic
浏览论文内容
中文总结 AI 辅助
SHAMS是一个针对黎凡特阿拉伯语的音频发音基准,包含1,300条平衡五种方言的话语,支持变音符号化、字素到音素转换等任务,用于评估LA语音技术进展。
中文摘要 AI 辅助
黎凡特阿拉伯语(LA)有数千万使用者,因此迫切需要共享基准来评估LA语音语言技术。鉴于LA的内部多样性及其不透明且非标准化的正字法,评估此类技术尤其具有挑战性。我们提出了SHAMS(沙姆标注多方言语音),这是一个包含1,300条话语的基准,这些话语来自开放音频语料库,在五种LA变体(城市和农村巴勒斯坦语,以及城市约旦语、黎巴嫩语和叙利亚语)之间保持平衡。每条话语在四个对齐层级中表示:音频、无元音正字法、带变音符号文本和音标转写。这种结构支持评估各种下游任务,如变音符号化、字素到音素转换、自动语音识别和音频到音素,这些任务以音频为基础并按变体分层。我们对开放和专有模型在这些任务上进行基准测试,以展示该基准在衡量LA进展方面的实用性。我们在以下网址发布SHAMS:此https URL。
英文摘要
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present SHAMS (SHami Annotated Multi-dialect Speech), a benchmark comprising 1,300 utterances drawn from open audio corpora, balanced across five LA varieties (Urban and Rural Palestinian, and Urban Jordanian, Lebanese, and Syrian). Each utterance is represented across four aligned tiers: audio, unvocalized orthography, diacritized text, and phonetic transcription. This structure supports evaluation of various downstream tasks such as diacritization, grapheme-to-phoneme conversion, automatic speech recognition, and audio-to-phoneme, grounded in audio and stratified by variety. We benchmark open and proprietary models across these tasks to demonstrate the utility of this benchmark for measuring progress across LA. We release SHAMS at https://shams-nlp.github.io .