arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIVID:面向越南NLP中比喻语言差距的文化基准测试集

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang

arXiv 2608.03095首次发表:更新:

AI 中文总结

研究团队构建了首个越南文化比喻语言基准VIVID,评估发现越南专用NLP模型远弱于多语言系统,顶尖模型得分不足满分50%,少样本提示未必提升性能,VIVID为相关研究提供关键工具。

AI 中文摘要

我们提出了VIVID(越南语习语用于验证和解释深度),这是首个用于评估越南语中基于文化的比喻语言理解的系统基准测试集。VIVID包含1636个习语和谚语,标注了五个复杂性特征(字面表达、语用细微差别、汉越词、罕见词汇、民间知识)以及七个语义主题。我们建立了结合生成式和判别式任务的评估框架,提出了一种基于方面提示的LLM-as-a-Judge方法,该方法经人工判断验证,Cohen's kappa值为0.792。对八个最先进模型的评估揭示了关键差距:越南专用模型的表现远低于多语言系统(VinaLLaMA-7B:0.13,而GPT-4o:2.46),即使是顶尖模型的得分也不到满分的50%。值得注意的是,少样本提示并不总能提升性能,GPT-4o因风格过拟合出现性能下降。我们的分析揭示了系统性失败,包括过度字面解读、词汇差距和语用扁平化,表明当前模型缺乏对细微比喻解读的文化能力。VIVID为推进文化丰富语境中的比喻语言理解提供了重要工具。

英文摘要

We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese. VIVID comprises 1,636 idioms and proverbs annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. We establish an evaluation framework combining generative and discriminative tasks, proposing an LLM-as-a-Judge approach with aspect-based prompting validated against human judgment (Cohen's kappa = 0.792). Evaluating eight state-of-the-art models reveals critical gaps: Vietnamese-specialized models drastically underperform multilingual systems (VinaLLaMA-7B: 0.13 vs. GPT-4o: 2.46), and even top models achieve less than 50% of maximum scores. Notably, few-shot prompting does not universally improve performance, with GPT-4o exhibiting degradation due to stylistic overfitting. Our analysis exposes systematic failures including literal over-interpretation, lexical gaps, and pragmatic flattening, demonstrating that current models lack cultural competence for nuanced figurative interpretation. VIVID provides an essential tool for advancing figurative language understanding in culturally rich contexts.

CommentsLREC 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑