arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越文化知识:评估大型语言模型的阿拉伯文化适切性

Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models

Enes Altinisik, Hamdy Mubarak, Masoomali Fatehkia, Husrev_Taha_Sencar Husrev Taha Sencar

arXiv 2609.16006首次发表:更新:

发表机构

Hamad Bin Khalifa University(哈马德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出AraBehave基准,评估LLM的阿拉伯文化适切性,发现其由规范性立场和扎实文化准确性组成,且立场易受提示影响,通用安全基准无法反映此问题。

AI 中文摘要

大型语言模型(LLMs)日益服务于那些期望受其文化背景影响的用户,然而大多数文化评估测试的是模型知道什么,而非其在给出开放式建议、观点和指导时如何表现。我们引入了AraBehave:包含1,623个植根于文化的阿拉伯语开放式提示,以及来自多个阿拉伯地区母语者的29,214条文化适切性判断,外加一个评分模型,其预测与未见系统上的人类判断高度相关(Pearson r=0.74)。评估三个以阿拉伯语为中心的和三个前沿的LLMs,我们发现文化适切性并非单一能力,而是分解为两个大致独立的组成部分:规范性立场和扎实的文化准确性。最佳通用模型和最佳阿拉伯语中心模型得分相同(3.84对3.83,满分5分),但几乎从不因相同原因失败:通用模型表现出扎实的事实基础,但规范性立场在文化上不恰当,因世俗框架和对文化已定论问题的虚假平衡而受到惩罚(占其低分理由的28-33%),而最佳阿拉伯语中心模型采取了预期立场,但因捏造的圣训和错误引用的经文而受到惩罚(占29%)。立场是廉价且脆弱的:一句文化指令使Gemini提升至4.57,超过所有阿拉伯语专业模型。相反,一个通用的“清晰客观地回答”提示使Allam-7B损失0.68分,而用英语提出相同问题则使除一个模型外的所有模型得分降低。扎实性反而与规模和阿拉伯语对齐数据相关,并且当以文化意识指令调优被文化中立语料库取代时,扎实性消失。通用安全基准看不到这些:它们在89以上饱和,而文化得分范围在2.71-3.84之间。我们将发布基准、注释和评分模型。

英文摘要

Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended recommendations, opinions, and guidance. We introduce AraBehave: 1,623 culturally grounded, open-ended Arabic prompts with 29,214 cultural-appropriateness judgments from native speakers across several Arab regions, plus a scoring model whose predictions correlate strongly with human judgments on unseen systems (Pearson r=0.74). Evaluating three Arabic-centric and three frontier LLMs, we find that cultural appropriateness is not a single capability but decomposes into two largely independent components: normative stance and grounded cultural accuracy. The best general-purpose and best Arabic-centric models score identically (3.84 vs. 3.83 of 5) yet almost never fail for the same reason: general-purpose models exhibit strong factual grounding but a culturally inappropriate normative stance, being penalized for secular framing and false balance on culturally settled matters (28--33% of their low-score rationales), while the best Arabic-centric model adopts the expected stance but is penalized for fabricated hadith and misquoted verses (29%). Stance is cheap and fragile: one sentence of cultural instruction lifts Gemini to 4.57, above every Arabic-specialized model. Conversely, a generic ``answer clearly and objectively'' prompt costs Allam-7B 0.68 points, while asking the same questions in English lowers scores for every model but one. Grounding instead tracks scale and Arabic alignment data, and disappears when culturally aware instruction tuning is replaced by a culture-neutral corpus. General safety benchmarks see none of this: they saturate above 89 while cultural scores span 2.71-3.84. We will release the benchmark, annotations, and the scoring model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑