CoCoA:面向大语言模型的上下文条件文化对齐
CoCoA: Context-Conditional Cultural Alignment for Large Language Models
浏览论文内容
中文总结 AI 辅助
CoCoA是面向大语言模型的上下文条件文化对齐框架,通过双上下文训练等方法,可降低实体文化偏见评分且对通用性能影响小,为缓解LLMs实体文化偏见提供新方向。
中文摘要 AI 辅助
大语言模型(LLMs)在不同文化语境中往往更倾向于与西方相关的实体。传统去偏方法追求统一的中立性,但文化偏见缓解需要上下文条件行为:当存在文化线索时偏好符合文化的实体,不存在时保持中立。我们提出CoCoA(Context-Conditional Cultural Alignment,上下文条件文化对齐)框架,通过对同一实体对在有和无文化线索的语境下进行双上下文训练来学习该行为。CoCoA结合对比对齐目标、校准与漂移正则化,通过目标感知梯度协调进行优化。我们在CAMeL和Camellia这两个以实体为中心的文化偏见基准上,针对10种语言设置和4种LLMs评估CoCoA。CoCoA将文化偏见评分从平均43降至24,同时保持近50.2的中立偏好,对5个标准基准的通用性能影响极小。这些发现表明,有效的文化对齐需要上下文条件建模而非统一去偏,为缓解LLMs中以实体为中心的文化偏见确立了新方向。
英文摘要
Large Language Models (LLMs) often favor Western-associated entities across cultural contexts. Conventional debiasing methods aim for uniform neutrality, but cultural bias mitigation demands context-conditional behavior, preferring culturally appropriate entities when cultural cues are present and remaining neutral when they are absent. We propose CoCoA (Context-Conditional Cultural Alignment), a framework that learns this behavior through dual-context training on the same entity pairs under contexts with and without cultural cues. CoCoA combines a contrastive alignment objective with calibration and drift regularization, optimized through goal-aware gradient reconciliation. We evaluate CoCoA on CAMeL and Camellia, two entity-centric cultural bias benchmarks, across ten language settings and four LLMs. CoCoA reduces the Cultural Bias Score from 43 to 24 on average while maintaining near-neutral preferences at 50.2, with minimal impact on general performance across five standard benchmarks. These findings highlight that effective cultural alignment requires context-conditional modeling rather than uniform debiasing, and establish a new direction for mitigating entity-centric cultural bias in LLMs.
发表机构
- Sungkyunkwan University(成均馆大学)
- Georgia Institute of Technology(佐治亚理工学院)
- University of Southern California(南加利福尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。