arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12341cs.CL

大语言模型在文化禁忌安全方面的“知识-行为差距”

The "Knowledge-Behavior Gap" in Cultural Taboo Safety of Large Language Models

Ying He, Sihang Jiang, Xingzhou Chen, Zhouhong Gu, Yiwei Gu, Minggui He, Shimin Tao, Hongxia Ma, Yanghua Xiao

首次发表
浏览论文内容

中文总结 AI 辅助

针对大语言模型文化禁忌安全评估的现有基准存在不足,本文推出首个公开基准CulShield,发现先进LLMs存在“知识-行为差距”,且语言语境变化会影响其文化禁忌安全。

中文摘要 AI 辅助

文化禁忌安全对部署大语言模型(LLMs)至关重要,因为文化上不敏感的输出可能会冒犯他人甚至造成社会危害。然而,现有的文化基准主要评估文化知识或价值观偏见,却忽略了LLMs能否识别并尊重文化禁忌,尤其是当禁忌隐含在看似无害的问题中时。此外,文化禁忌具有隐含性和语境依赖性,因此对可靠评估构成了独特挑战。为解决这些差距,我们推出了首个专门用于评估和提升LLMs文化禁忌安全的公开基准——CulShield。CulShield覆盖77个国家和地区,包含2020多个禁忌,从显性知识和隐性行为两个维度评估模型。对多个先进LLMs(如GPT-4o-mini、Gemini-2.5-pro)的实验显示存在明显的“知识-行为差距”:模型在交互过程中常常无法应用已知的禁忌。我们进一步表明,语言语境的变化会显著影响LLMs的文化禁忌安全。代码和数据可在此处获取:this https URL。

英文摘要

Cultural taboo safety is essential for deploying large language models (LLMs), as culturally insensitive outputs may cause offense or even social harm. However, existing cultural benchmarks primarily assess cultural knowledge or values biases, while overlooking whether LLMs can recognize and respect cultural taboos, especially when taboos are implicitly hidden in seemingly harmless questions. Besides, cultural taboos are implicit, and context-dependent, thus poss unique challenges for reliable evaluation. To address these gaps, we introduce \textbf{CulShield}, the first public benchmark dedicated to evaluating and improving the cultural taboo safety of LLMs. CulShield spans 77 countries and territories, and includes over 2,020 taboos. It evaluates models along both explicit knowledge and implicit behaviors. Experiments on several advanced LLMs (e.g., GPT-4o-mini, Gemini-2.5-pro) reveal a clear ``knowledge-behavior gap'': models often fail to apply known taboos during interaction. We further show that variations in linguistic context can significantly affect LLMs' cultural taboo safety. Code and data is accessible here: https://github.com/hedyHe/CulShield.

发表机构

  • College of Computer Science and Artificial Intelligence, Fudan University(复旦大学计算机与人工智能学院)
  • Huawei(华为)

机构由 AI 辅助整理,请以论文原文为准。

↑