arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AraSSM:面向阿拉伯语掩码语言建模的双向状态空间编码器

AraSSM: A bidirectional state-space encoder for Arabic masked language modeling

Ahmed Amine Aliane, Hassina Aliane, Nasredine Semmar

arXiv 2608.08256首次发表:更新:

发表机构

Arabic Institute for Translation; CERIST; CEA LIST(阿拉伯翻译学院; CERIST(阿尔及利亚科学技术研究中心); CEA LIST(法国原子能和替代能源委员会电子信息技术与实验室))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出面向阿拉伯语的双向状态空间编码器AraSSM,经掩码语言建模预训练后,在四项阿拉伯语NLU基准上的表现接近或优于同规模Transformer,验证了消费级硬件训练的有效性。

AI 中文摘要

AraBERT、MARBERT和CAMeLBERT等预训练Transformer编码器已成为阿拉伯语自然语言理解的标准骨干,但它们的自注意力机制随序列长度呈二次方缩放,限制了长文档处理的效率。Mamba作为一种选择性状态空间模型(SSM),提供线性时间序列建模能力,是注意力机制的竞争性替代方案,但目前尚无专门针对阿拉伯语预训练的双向Mamba编码器。本文提出AraSSM,这是一种通过掩码语言建模在阿拉伯语维基百科与CulturaX文本结合的语料库上预训练的双向Mamba编码器,在4块消费级NVIDIA RTX 2080Ti GPU(11GB显存)上进行端到端训练,耗时约10天。本文遵循AraBERT提出的每任务评估协议,通过在四个成熟的阿拉伯语NLU基准上微调来评估AraSSM,这些基准涵盖情感分类(HARD)、命名实体识别(ANERcorp)、抽取式问答(ARCD)和自然语言推理(XNLI-ar),结果以三次微调种子的均值±标准差形式呈现。AraSSM在情感分类上达到或超过已发表的基础规模Transformer基准(在HARD上准确率为96.37±0.03%),在抽取式问答(ARCD上EM为32.19±1.07,F1为63.79±0.25)和命名实体识别(ANERcorp上实体级F1为81.54±0.30)上具有竞争力,在自然语言推理(XNLI-ar上准确率为72.83±0.07%)上略低于基础规模Transformer的表现范围,尽管它完全在消费级硬件上从头训练,而非大规模加速器集群。

英文摘要

Pretrained Transformer encoders such as AraBERT, MARBERT, and CAMeLBERT have become the standard backbone for Arabic natural language understanding, but their self-attention mechanism scales quadratically with sequence length, which limits efficiency on long documents. Mamba, a selective state-space model (SSM), offers linear-time sequence modeling as a competitive alternative to attention, yet no dedicated bidirectional Mamba encoder pretrained specifically for Arabic currently exists. We introduce AraSSM, a bidirectional Mamba encoder pretrained via masked language modeling on a corpus combining Arabic Wikipedia and CulturaX text, trained end-to-end on four consumer-grade NVIDIA RTX 2080Ti GPUs (11GB) over approximately ten days. We evaluate AraSSM by fine-tuning on four established Arabic NLU benchmarks covering sentiment classification (HARD), named entity recognition (ANERcorp), extractive question answering (ARCD), and natural language inference (XNLI-ar), following the per-task evaluation protocol introduced by AraBERT, and report results as mean +/- standard deviation across three fine-tuning seeds. AraSSM matches or exceeds published base-sized Transformer baselines on sentiment classification (96.37 +/- 0.03% accuracy on HARD), is competitive on extractive QA (32.19 +/- 1.07 EM, 63.79 +/- 0.25 F1 on ARCD) and named entity recognition (81.54 +/- 0.30 entity-level F1 on ANERcorp), and trails the base-sized Transformer range on natural language inference (72.83 +/- 0.07% accuracy on XNLI-ar), despite being trained entirely from scratch on consumer hardware rather than large-scale accelerator clusters.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑