arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当一个名字并非名字时:低资源语言模型中文化纠缠的孟加拉语同形异义词的基准数据集和提炼推理

When a Name Is Not a Name: A Benchmark Dataset and Distilled Reasoning for Culturally Entangled Bangla Homographs in Low-Resource LLMs

Md. Asaduzzaman Shuvo

arXiv 2607.17828首次发表:更新:

发表机构

United International University(联合国际大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对低资源语言模型中孟加拉语同形异义词文化纠缠问题,构建基准数据集。发现模型有主导意义偏差,特定模型失败。提出对比性思维链提示及提炼文化解释,能减少偏差,让小模型正确推理,提升系统性能。

AI 中文摘要

许多孟加拉语单词既是人名又是具有文化内涵的普通名词,如“Maya”既是女孩名字又是深情同情之意。选择正确释义需要文化知识,而现代语言模型预训练数据中此类知识稀缺。我们引入文化纠缠同形异义词(CEH)消歧并构建了一个包含1516个经专家验证句子(3032个标注出现情况)的孟加拉语基准,其中一个单词两次出现有两种不同释义,并标注了基于文化的类别及推理依据。研究发现模型存在主导意义偏差,特定孟加拉语模型在各种提示机制下均失败,而对比性思维链提示可大幅减少偏差,提炼文化解释能让小模型正确推理而非记忆标签,将主导意义偏差从高达100%降至5%以下,使失败的特定孟加拉语模型成为最强系统。数据集和代码可通过链接获取。

英文摘要

Many Bangla words are at once personal names and culturally loaded common nouns, "Maya" is both a girl's name and a word for affectionate compassion. Choosing the right reading demands cultural knowledge that is scarce in the pretraining data of modern language models. We introduce Culturally Entangled Homograph (CEH) disambiguation and build a Bangla benchmark of 1,516 expert-verified sentences (3,032 labelled occurrences) in which one word appears twice with two distinct readings, each labelled with a culturally grounded category and an explanation of the reasoning behind it. Across open- and closed-source models, we find a systematic dominant-meaning bias: models default to the common-noun sense and overlook the name. A Bangla-specific model fails under every prompting regime we test, showing that language-specific pretraining alone does not confer cultural grounding. We further show that contrastive chain-of-thought prompting can sharply reduce this bias without training, and that distilling cultural explanations teaches small (1-3B) models to reason toward the correct reading rather than memorise labels, cutting dominant-meaning bias from as high as 100% to under 5% and turning the failed Bangla-specific model into our strongest system. Dataset and code are available at https://github.com/ashuvo25/BanglaCEH.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑