发表机构
University of Washington; Centific(华盛顿大学; Centific)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM行为评估忽视文化背景的问题,提出CuBEs文化情境行为评估框架,通过注入文化情境并构建12种文化数据集,揭示标准评估无法捕捉的跨文化行为差异,强调全球部署需文化情境化测试。
AI 中文摘要
评估大语言模型(LLM)行为(如谄媚、自我偏好或过度自信)的发生及其触发因素,对于预测模型在现实世界部署中的风险至关重要。然而,现有的情境化行为评估通常忽略文化背景,限制了其在日益全球化的用户群体中的泛化能力。为弥补这一空白,我们提出了CuBEs——文化情境行为评估,旨在探测不同用户文化背景下的响应模式。我们首先扩展了一个自动化测试流程,将文化背景注入行为测试场景及其后续评估中。通过构建一个涵盖12种不同文化、捕捉行为理解细微维度的人工标注数据集,我们评估了该流程的文化适应性。我们的数据集揭示了显著的跨文化差异,而这些差异是“一刀切”式判断所无法捕捉的。通过评估13个开源和闭源LLM,我们发现,在评估场景中引入文化情境性会显著改变行为出现的差异。例如,我们的基线实验测试政治偏见时捕捉到的是西方本地化的政治维度,如美国保守派与进步派的分歧,而非西方文化情境下的评估则浮现出完全不同的偏见轴,如宗教和殖民政治问题。我们的研究结果表明,标准的、文化无关的评估无法捕捉这些变化,凸显了全球部署中文化情境行为测试的必要性。
英文摘要
Evaluating the occurrence and triggers of large language model (LLM) behaviors - such as sycophancy, self-preference, or over-confidence - is critical for predicting real-world model deployment risks. However, existing situated behavioral evaluations typically ignore cultural context, limiting their generalizability across an increasingly global user base. To address this gap, we propose CuBEs - Culturally-situated Behavior Evaluations that probe for response patterns across diverse user cultures. We first extend an automated testing pipeline to inject cultural context into behavioral test scenarios and subsequent evaluation. We assess the cultural adaptability of this pipeline by building a human-labeled dataset that captures nuanced dimensions of behavior understanding across 12 distinct cultures. Our dataset reveals significant cross-cultural variations that one-size-fits all judgments fail to capture. Through evaluating 13 open- and closed-source LLMs, we find that introducing cultural situatedness in the evaluation scenario creates significant variation in the presence of a behavior. For example, while our baseline experiments testing for political bias capture localized Western political dimensions like the American conservative-progressive divide, non-Western culturally situated evaluations surface entirely different axes of bias such as religious and colonial political issues. Our findings demonstrate that standard, culturally-agnostic evaluations fail to capture these shifts, highlighting the necessity of culturally situated behavioral testing for global deployments.