arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CIBuzzBench:中文网络流行语跨语言理解基准

CIBuzzBench: A Benchmark for Cross-Lingual Understanding of Chinese Internet Buzzwords

Yifan Wang, Junyu Lu, Qifan Wang, Shun Zhang, Chaozhuo Li, Jiahao Liu, Zhijun Cao, Lingbin Bu, Fanliang Bu

arXiv 2609.21722首次发表:更新:

发表机构

Dalian University of Technology; Meta AI; Beijing University of Posts and Telecommunications; Meituan; Beijing Police College(大连理工大学; Meta AI; 北京邮电大学; 美团; 北京警察学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对大语言模型跨语言理解中文网络流行语能力不足的问题,构建首个中译英基准CIBuzzBench,包含3001个流行语及多维度标注,设计含义解释、对应匹配和危害检测三项任务,实验表明现有模型在非字面解释、稳健匹配和危害检测上仍存在显著困难。

AI 中文摘要

中国社交媒体产生了大量且不断演变的网络流行语词汇,其含义往往是非字面的,并深深植根于当地文化和语用语境中。现有研究主要集中于在中文语境下解读这些流行语,而很少探究大语言模型(LLMs)能否将这种基于文化的知识跨语言迁移,并准确地在英语中传达其预期含义。这种跨语言能力对于安全性也至关重要,因为有害表达可能通过文化特有的谐音、委婉语、反讽或编码语言来掩盖其冒犯性内容。在本文中,我们研究了先进大语言模型跨语言理解中文网络流行语的能力。为此,我们引入了CIBuzzBench,这是首个用于中文网络流行语跨语言(中译英)理解的基准。CIBuzzBench包含3,001个中文网络流行语,并标注了英文含义解释、英文对应词、类别标签和有害性标签。基于这些标注,我们设计了三个评估任务:含义解释、跨语言对应匹配和基于文化的危害性检测。我们在英文提示和中文提示两种设置下评估了具有代表性的最先进的专有模型和中文大语言模型。我们的结果表明,大语言模型在中文网络流行语的跨语言理解方面仍然存在困难,尤其是在细粒度的非字面解释、选项扰动下的稳健对应匹配以及校准的危害性检测方面。这些发现凸显了基于文化的语言现象对多语言大语言模型和面向安全的评估所带来的持续挑战。数据集和代码可在以下网址获取:https URL。

英文摘要

Chinese social media has generated a vast and continually evolving lexicon of internet buzzwords whose meanings are often non-literal and deeply rooted in local cultural and pragmatic contexts. Existing research has primarily focused on interpreting these buzzwords within Chinese, leaving largely unexplored whether LLMs can transfer such culturally grounded knowledge across languages and accurately convey the intended meanings in English. This cross-lingual capability is also critical for safety, as harmful expressions may obscure their offensive content through culture-specific homophony, euphemism, irony, or coded language. In this paper, we investigate the ability of advanced LLMs to understand Chinese internet buzzwords across languages. To this end, we introduce CIBuzzBench, the first benchmark for cross-lingual Chinese-to-English understanding of Chinese internet buzzwords. CIBuzzBench comprises 3,001 Chinese internet buzzwords annotated with English meaning explanations, English equivalents, category labels, and harmfulness labels. Based on these annotations, we design three evaluation tasks: Meaning Explanation, Cross-lingual Equivalent Matching, and Culturally Grounded Harmfulness Detection. We evaluate representative state-of-the-art proprietary and Chinese LLMs under both English- and Chinese-prompting settings. Our results show that LLMs continue to struggle with the cross-lingual understanding of Chinese internet buzzwords, particularly in fine-grained non-literal interpretation, robust equivalent matching under option perturbations, and calibrated harmfulness detection. These findings highlight the persistent challenges posed by culturally grounded language phenomena for multilingual LLMs and safety-oriented evaluation. The dataset and code are available at https://github.com/SuperYFan/CIBuzzBench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑