arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.10758cs.CLcs.AIcs.LG

名不副实的多语言?LLMs在乌尔都语中的文化与语言弱点

Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu

  • University of Konstanz(康斯坦茨大学)
  • University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
  • Monash University Malaysia(莫纳什大学马来西亚校区)
  • Dalhousie University(达尔豪斯大学)

机构由 AI 辅助整理,请以论文原文为准。

Farah Adeeba, Abdul Rafae Khan, Rajesh Bhatt, Hassan Sajjad

AI总结:

本研究通过生成93个乌尔都语故事并人工标注错误,发现多语言LLMs在低资源语言中存在基本语法语义错误、缺乏连贯性及文化浅薄等问题,且少样本提示无法解决文化错误。

AI中文摘要:

多语言大语言模型(LLMs)越来越多地被用于开放式文本生成,然而它们在低资源语言中的行为仍然知之甚少。在这项工作中,我们质疑多语言LLMs在故事生成任务中的生成正确性和可靠性。我们将乌尔都语视为一种具有代表性的低资源语言。我们生成了乌尔都语故事集(Urdu-Stories),这是一个包含93个故事的语料库,这些故事由三个当代LLMs(GPT-5.1、Qwen-3-Max、DeepSeek-3.1)生成。我们手动标注了其中存在的错误,采用一个包含九个标签的语言、语义和文化分类体系。我们的显著发现表明,LLMs经常犯基本的语法和语义错误。这些故事缺乏连贯性,存在不自然的重复,并表现出普遍的文化浅薄。我们进一步通过少样本提示(few-shot prompting)表明,文化和上下文错误在很大程度上仍未得到解决。我们的发现凸显了当前LLMs作为低资源语言内容生成和信息检索可靠来源的局限性。

英文摘要:

Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.

↑