发表机构
Kobe University; University of Hyogo(神户大学; 兵库县立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究评估日本大语言模型对五种日语特定拼写错误的鲁棒性,发现字符换位和替换显著降低性能,而输入法相关错误影响较小,凸显了实际输入环境中鲁棒性评估的重要性。
AI 中文摘要
大语言模型(LLMs)在各种自然语言处理任务中已展现出强大的性能。然而,它们对拼写错误的鲁棒性仍未得到充分探索,尤其是在日语中,文本输入涉及多种书写系统和基于输入法的转换。在本研究中,我们评估了日本大语言模型对现实日语特定拼写错误的鲁棒性。我们引入了五种拼写错误类别:字符换位、字符替换、同音字转换、日语输入法转换和全角转换。这些扰动应用于三个日语基准数据集(JMMLU、JCommonsenseQA和JamC-QA),并评估了十一个日语和多语言大语言模型。结果表明,字符换位和字符替换拼写错误在各基准上持续降低准确率,而输入法转换、全角转换和同音字转换的影响相对有限。这些发现揭示了当前日本大语言模型对现实日语输入错误仍然脆弱,尤其是那些大幅扭曲原始输入的错误,凸显了在实际输入环境中进行鲁棒性评估的重要性。
英文摘要
Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this study, we evaluate the robustness of Japanese LLMs against realistic Japanese-specific typos. We introduce five typo categories: Character Transposition, Character Replacement, Homophone Conversion, Japanese IME Conversion, and Full-Width Conversion. These perturbations are applied to three Japanese benchmark datasets (JMMLU, JCommonsenseQA, and JamC-QA), and eleven Japanese and multilingual LLMs are evaluated. The results show that Character Transposition and Character Replacement typos consistently reduce accuracy across benchmarks, whereas IME Conversion, Full-Width Conversion, and Homophone Conversion have relatively limited impact. These findings reveal that current Japanese LLMs remain vulnerable to realistic Japanese typing errors, particularly those that substantially distort the original input, highlighting the importance of robustness evaluation in practical input environments.