发表机构
University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对语言模型用于跨语言文化道德决策时存在的多语言性忽视问题,提出MET两步提示法及MET-D自我蒸馏训练,引入MCLASH基准,实验表明MET-D提升模型性能,揭示不同文化有益依据差异,为多语言道德推理开辟道路。
AI 中文摘要
语言模型越来越多地用于跨语言和文化背景的道德决策,但现有工作在三个方面忽视了多语言性:多语言评估基准使用直接翻译,未能适应特定文化项目;道德推理的推理时间方法依赖于静态的、以英语为中心的框架,缺乏道德理论基础;道德决策的训练方法通常需要来自更强模型或人类注释者的昂贵监督。我们通过三项贡献解决了这些差距。首先,我们引入了MCLASH,这是一个多语言道德决策基准,以捕捉跨语言的文化情境化道德直觉和社会规范。其次,我们提出了MET(基于理论推理的多语言伦理学),这是一种两步提示方法,基于从心理学和哲学中提取的专家策划、基于理论的依据:模型首先选择特定情境和文化的依据,然后用用户的母语对其进行推理。第三,我们引入了MET-D(MET蒸馏),它通过一个不需要外部监督的自我蒸馏训练阶段来增强第二步。MET-D在不同大小和家族的所有三个模型(Qwen3-4B、Qwen3-8B、Gemma3-4B)上的宏观F1均优于基础模型,在MCLASH上平均提高3.71分,在MMoralExceptQA上提高4.23分,Qwen3-8B上马来语的MCLASH增益峰值为12.94分。我们进一步揭示,MET-D平均增加了62.13分的母语推理,并且有益的依据在不同文化中系统地不同。总之,这些贡献为文化对齐、基于理论的多语言道德推理开辟了道路。
英文摘要
Language models are increasingly used for moral decision-making across diverse linguistic and cultural contexts, yet existing work overlooks multilinguality on three aspects: 1) multilingual evaluation benchmarks use direct translation, failing to adapt culture-specific items; 2) inference-time methods for moral reasoning rely on static, English-centric scaffolds and lack grounding in moral theory; 3) training methods for moral decision-making typically require expensive supervision from stronger models or human annotators. We address these gaps with three contributions. First, we introduce MCLASH, a multilingual moral decision-making benchmark to capture culturally situated moral intuitions and social norms across languages. Second, we propose MET (Multilingual Ethics with Theory-grounded reasoning), a two-step prompting method built on expert-curated, theory-based grounds drawn from psychology and philosophy: the model first selects situation- and culture-specific grounds, then reasons over them in the native language of the user. Third, we introduce MET-D (MET-Distillation), which enhances the second step through a self-distillation training stage that requires no external supervision. MET-D improves macro-F1 over the base model on all three models of different sizes and families (Qwen3-4B, Qwen3-8B, Gemma3-4B), by an average of 3.71 points on MCLASH and 4.23 on MMoralExceptQA, with a peak MCLASH gain of 12.94 points for Malay on Qwen3-8B. We further reveal that MET-D increases native-language reasoning by 62.13 points on average, and that beneficial grounds differ systematically across cultures. Together, these contributions open the path for culture-aligned, theory-grounded multilingual moral reasoning.
CommentsPublished as a conference paper at COLM 2026