发表机构
Technical University Dortmund; Research Center Trustworthy Data Science and Security; BRAC University; United International University(多特蒙德工业大学; 可信数据科学与安全研究中心; BRAC大学; 联合国际大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建了首个孟加拉语习语基准数据集及配套MCQ数据集,通过多任务评估发现不同LLM在习语相关任务上表现有差异,各有优势,成果为低资源语言习语理解研究提供了资源。
AI 中文摘要
习语表达是自然语言的重要组成部分,反映了文化细微差别,对计算模型构成独特挑战,尤其是在低资源语言领域。本文中,我们推出首个大规模孟加拉语习语基准数据集,并补充了用于习语含义识别的合成多项选择题(MCQ)数据集。我们针对三项与习语相关的任务(释义、习语跨度检测、含义识别),结合零样本和少样本提示策略,对近期大型语言模型(LLMs)开展全面评估。结果显示模型性能存在显著差异,没有任何一款LLM在所有任务中始终优于其他模型。值得注意的是,Phi-4-mini-instruct在释义任务中表现出色,Kimi-K2-32b-instruct在跨度检测任务中表现出色,Gemini-2.5-flash在含义识别任务中表现出色。我们认为,所提出的数据集和分析将为未来研究提供宝贵资源,以提升LLM对习语表达的理解,尤其是针对孟加拉语及其他低资源语言的相关研究。
英文摘要
Idiomatic expressions are an integral part of natural language, reflecting cultural nuances and posing unique challenges for computational models, particularly in low-resource languages. In this paper, we present the first large-scale benchmark dataset of Bangla idioms, complemented by a synthetic multiple-choice question (MCQ) dataset for idiom meaning identification. We conduct a comprehensive evaluation of recent large language models (LLMs) across three idiom-related tasks: paraphrasing, idiom span detection, and meaning identification, leveraging zero-shot and few-shot prompting strategies. Our results reveal substantial variability in model performance, with no single LLM consistently outperforming others across all tasks. Notably, Phi-4-mini-instruct excels in paraphrasing, Kimi-K2-32b-instruct in span detection, and Gemini-2.5-flash in meaning identification. We believe that our datasets and analyses will provide valuable resources to guide future research in improving LLM comprehension of idiomatic expressions, particularly in Bangla and other low-resource languages.