大型语言模型对越南方言的鲁棒性如何?
How Robust Are LLMs to Vietnamese Dialects?
浏览论文内容
中文总结 AI 辅助
本研究通过构建VialectBench基准,评估指令微调LLM对越南方言的鲁棒性,发现方言输入会导致模型性能下降,且不同方言组的影响存在显著差异,标准越南语上的性能不代表方言场景下的可靠性。
中文摘要 AI 辅助
大型语言模型(LLMs)通常在标准书面越南语上进行评估,但日常交流中常使用保留语义但表层形式不同的地区方言。现有越南方言研究多通过方言到标准语的归一化解决该问题,而非测量模型在越南方言输入下的失效情况。为填补这一空白,我们首次系统评估了LLM对越南方言变异在多项任务中的鲁棒性,量化性能下降与失效模式。我们推出VialectBench(越南方言基准测试),这是一个受控基准,用于测试模型决策在六个越南方言组间是否保持稳定。VialectBench包含400个标准越南语源实例和2400个人工撰写的方言改写文本,覆盖情感识别(ER)、自然语言推理(NLI)、问答(QA)、多项选择问答(MCQA)任务。对固定参考语言模型的数据集评估显示,方言改写会引发可测量的模型相对似然偏移,且长度与标准越南语对应文本几乎相等。在十个指令微调模型中,方言输入使平均性能下降2.82%,无评估模型完全具备方言不变性。四个任务均受影响,其中QA的平均下降幅度最大。鲁棒性在不同方言组间差异显著:PNT3和PNT2分别导致最大平均性能下降6.17%和4.73%,而PNB使平均性能略有提升0.42%。中部方言组(PNT1-PNT4)在所有模型中产生最高的有害翻转率,达6.54%。这些发现表明,标准越南语上的优异表现无法保证在保留语义的地区变异下行为可靠。
英文摘要
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalization instead of measuring how the model fails under Vietnamese dialectal inputs. To address this gap, we present the first systematic evaluation of LLM robustness to Vietnamese dialect variation across multiple tasks, quantifying performance degradation and failure patterns. We introduce VialectBench (Vietnamese Dialects Benchmarking), a controlled benchmark for testing whether model decisions remain stable across six Vietnamese dialect groups. VialectBench contains 400 Standard Vietnamese source instances and 2,400 human-written dialectal rewrites spanning emotion recognition (ER), natural language inference (NLI), question answering (QA), and multiple-choice question answering (MCQA). Dataset evaluation with a fixed reference language model shows that the dialectal rewrites induce a measurable model-relative likelihood shift while remaining nearly equal in length to their Standard counterparts. Across ten instruction-tuned models, dialectal inputs reduce average performance by 2.82%, and no evaluated model is fully dialect-invariant. All four tasks are affected, with QA showing the largest average degradation. Robustness also varies substantially across dialect groups: PNT3 and PNT2 cause the largest average performance drops, at 6.17% and 4.73%, respectively, whereas PNB slightly improves average performance by 0.42%. The Central dialect group (PNT1-PNT4) also yields the highest average harmful-flip rate across all models, at 6.54%. These findings show that strong performance on Standard Vietnamese does not guarantee reliable behavior under meaning-preserving regional variation.