arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01788cs.CLcs.AI

VakyArth:评估跨印度语言的大语言模型语用能力

VakyArth: Evaluating Pragmatic Competence in LLMs across Indic Languages

Usneek Singh, Poorvaja Veera Balaji Kumar, Parth Nanda, Anand Madhusoodanan, Geyang Guo, Wei Xu, Junyi Jessy Li

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出首个印度语言语用基准VakyArth,评估多语言LLM在印度语言文化相关语用现象上的表现,发现其存在系统性语用推理失败及语言间差异。

中文摘要 AI 辅助

现实世界的交流常常需要语用推理:解读通过语境和文化习俗隐含的意义,而非字面表述的含义。现有的语用评估大多局限于英语和高资源语言,尽管印度语言具有语言和文化多样性,但仍未被探索。我们推出VakyArth,这是首个针对印度语言的语用基准,设计为涵盖印地语、旁遮普语、泰米尔语和马拉雅拉姆语的诊断评估。VakyArth通过多项选择题、自然语言推理和翻译任务评估模型在五种语用现象上的表现:指示语、言语行为、隐涵、社会语用学和连贯性,所有题目均由母语者编写。在不同家族和规模的多语言大语言模型(LLM)中,我们发现其在源于印度语言和文化习俗的语用意义上存在持续失败。我们的分析显示,不同语言和任务间存在系统性差异:在所有模型-语言组合中,多项选择题准确率超过自然语言推理准确率;翻译表现无法可靠反映语用理解;印度-雅利安语系语言在翻译上优于达罗毗荼语系语言。我们进一步表明,自动翻译指标可能忽略流畅但语用不忠实的输出,尤其是在隐涵和指示语方面。

英文摘要

Real-world communication often requires pragmatic reasoning: interpreting meanings implied through context and cultural convention rather than stated literally. Existing pragmatic evaluation remains largely limited to English and high-resource languages, leaving Indic languages unexplored despite their linguistic and cultural diversity. We introduce VakyArth, the first pragmatic benchmark for Indic languages, designed as a diagnostic evaluation covering Hindi, Punjabi, Tamil, and Malayalam. VakyArth evaluates models across five phenomena: deixis, speech acts, implicature, social pragmatics, and coherence; through multiple-choice questions, natural language inference, and translation, with all items authored by native speakers. Across multilingual large language models (LLMs) of varying families and sizes, we find consistent failures on pragmatic meanings rooted in Indic linguistic and cultural conventions. Our analysis shows systematic differences across languages and tasks: MCQ accuracy exceeds NLI accuracy in all model-language combinations, translation performance does not reliably track pragmatic understanding, and Indo-Aryan languages show a translation advantage over Dravidian languages. We further show that automatic translation metrics can miss fluent but pragmatically unfaithful outputs, especially for implicature and deixis.

发表机构

  • Georgia Institute of Technology(佐治亚理工学院)
  • University of Texas at Austin(德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑