arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05510cs.CL

不同扰动,不同机制:理解用于零样本方言鲁棒性的持续预训练

Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness

Aarohi Srivastava, David Chiang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究针对大型语言模型,系统对比6种训练条件下的6种扰动持续预训练策略,发现字符噪声CPT可提升零样本方言鲁棒性且保留标准性能,不同策略的鲁棒机制存在差异,为相关策略选择提供指导。

中文摘要 AI 辅助

方言变异仍是多语言语言模型面临的主要挑战。基于扰动的持续预训练(CPT)已成为提升鲁棒性的有前景方法,但现有研究大多孤立评估单个扰动策略,且对其奏效原因的见解有限。我们针对大型语言模型(LLM)中用于多语言方言鲁棒性的基于扰动的CPT开展系统研究,在9个德语、意大利语和阿拉伯语方言任务中对比6种训练条件。基于扰动的CPT,尤其是字符噪声CPT,可持续提升零样本方言鲁棒性,同时在很大程度上保留标准变体的性能。更重要的是,我们表明,具有相似下游性能的方法会引发截然不同的鲁棒性机制,呈现出不同的语言模型适应、表征对齐及预测修复模式。我们的结果为合成表层变异如何提升鲁棒性提供了更完整的理解,并为在多语言和方言场景中选择CPT策略提供了实用指导。

英文摘要

Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work. We present a systematic study of perturbation-based CPT for multilingual dialect robustness in LLMs, comparing six training conditions across nine German, Italian, and Arabic dialect tasks. Perturbation-based CPT, especially character-noised CPT, consistently improves zero-shot dialect robustness while largely preserving standard variety performance. More importantly, we show that methods with similar downstream performance induce distinct mechanisms of robustness, exhibiting different patterns of language model adaptation, representational alignment, and prediction repair. Our results provide a more complete understanding of how synthetic surface variation improves robustness and offer practical guidance for selecting CPT strategies in multilingual and dialectal settings.

发表机构

  • University of Notre Dame(圣母大学)

机构由 AI 辅助整理,请以论文原文为准。

↑