arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

阿拉伯语语言模型中的对抗鲁棒性评估

Evaluation of Adversarial Robustness in Arabic Language Models

Anwar Alajmi, Ayed Salman, Imtiaz Ahmad

arXiv 2607.25814首次发表:更新:

发表机构

Kuwait University; College of Business Studies, Public Authority of Applied Education and Training(科威特大学; 应用教育与培训公共管理局商业研究学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究评估五个先进阿拉伯语语言模型在不同攻击下的鲁棒性,探索对抗训练防御技术,发现不同攻击对模型准确率影响各异,对抗训练提升了整体弹性,但面对字符级噪声仍有挑战,凸显了当前防御策略在阿拉伯语中的潜力与局限。

AI 中文摘要

近期阿拉伯语语言模型展现出卓越能力,但也暴露出漏洞,对抗攻击是主要安全风险之一。本研究旨在评估五个先进阿拉伯语语言模型在不同粒度级别和不同示例生成策略下的一组独特阿拉伯语对抗攻击中的鲁棒性,并探索基于对抗训练的防御技术。结果表明,添加变音符可使某些模型准确率降低92%,词级攻击中操纵阿拉伯连词会导致准确率下降高达58%,句子级攻击中释义可使受害模型性能平均降低76%。对抗训练虽提高了整体弹性,但仍存在挑战。这些发现凸显了当前防御策略在阿拉伯语这种形态丰富语言中的潜力和局限性。

英文摘要

The emergence of the recent outstanding capabilities of Arabic Language Models has opened doors for exposing their vulnerabilities. One of the major security risks associated with such Natural Language Processing models is adversarial attacks. These attacks can deceive the model into the wrong prediction, raising critical model security and safety concerns. This study aims to assess the robustness of five state-of-the-art Arabic Language Models under a distinct set of Arabic adversarial attacks applied at various levels of granularity and using different example generation strategies. We also explore a defense technique based on adversarial training to enhance model robustness. The results show that insertion of diacritics can reduce the accuracy of some models by 92% while maintaining a low perturbation distance. For word-level attacks, manipulating Arabic conjunctions preserves high semantic similarity scores, low perturbation distance, and leads to an accuracy degradation of up to 58%. For sentence-level attacks, paraphrasing proves its effectiveness by an average reduction of 76% in the victim models' performance. While adversarial training improves overall resilience, with MARBERT being the most robust and AraBERT showing the greatest relative gains, challenges persist, particularly against character-level noise. These findings highlight both the potential and limitations of current defense strategies in morphologically rich languages like Arabic.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑