CRISP:面向隐含情境得体性的文化奖励建模
CRISP: Cultural Reward Modeling for Implicit Situated Propriety
浏览论文内容
中文总结 AI 辅助
本文提出文化情境奖励模型CRISP-RM及规范基础监督(NGS),用于开放式社交场景中的文化得体性判断,并通过协作多智能体框架构建NormCompass测试平台,实验证明其在奖励建模和策略优化上均优于通用模型。
中文摘要 AI 辅助
随着大语言模型(LLMs)在各国和各地区的部署日益增多,识别并恰当回应多元文化背景的能力变得越来越重要。然而,现有研究主要聚焦于文化知识或具有预定义响应空间的任务,而开放式文化情境行为仍相对未被充分探索。在这项工作中,我们引入了CRISP-RM,一种文化情境奖励模型,它根据开放式社交场景中的文化得体性来分配奖励。在策略优化过程中,我们进一步引入了规范基础监督(NGS),提供指导以增强策略对相关文化规范的敏感性。为了构建文化情境数据,我们采用了一个协作式多智能体框架,将隐含的文化规范实例化为多样的社交场景,并进一步整理出NormCompass作为专用测试平台。我们进行了全面的实验,以评估CRISP-RM在奖励建模和策略优化两方面的有效性。Best-of-\(N\)实验表明,CRISP-RM持续优于强大的通用奖励模型。在GRPO策略优化过程中,CRISP-RM总体上改善了文化情境行为,而结合NGS则带来了进一步的提升。进一步的分析证明了CRISP-RM在区分文化得体行为方面具有优势,超越了表面的流畅性和礼貌性,而NGS在策略优化过程中通过改进规范基础提供了补充性的增益。
英文摘要
As large language models (LLMs) are increasingly deployed across countries and regions, the ability to recognize and respond appropriately to diverse cultural contexts becomes increasingly important. However, existing research has largely focused on cultural knowledge or tasks with predefined response spaces, while open-ended culturally situated behavior remains comparatively underexplored. In this work, we introduce CRISP-RM, a culturally situated reward model that assigns rewards according to cultural appropriateness in open-ended social scenarios. During policy optimization, we further introduce Norm Grounding Supervision (NGS), providing guidance that enhances the policy's sensitivity to relevant cultural norms. To construct culturally situated data, we employ a collaborative multi-agent framework that instantiates implicit cultural norms into diverse social scenarios and further curate NormCompass as a dedicated testbed. We conduct comprehensive experiments to evaluate the effectiveness of CRISP-RM in both reward modeling and policy optimization. Best-of-\(N\) experiments show that CRISP-RM consistently outperforms strong general reward models. During GRPO policy optimization, CRISP-RM generally improves culturally situated behavior, while incorporating NGS yields further gains. Further analyses demonstrate the advantages of CRISP-RM in distinguishing culturally appropriate behavior beyond superficial fluency and politeness, while NGS provides complementary gains during policy optimization by improving norm grounding.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- Huawei Technologies Co., Ltd(华为技术有限公司)
- Peng Cheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。