发表机构
Mohamed bin Zayed University of Artificial Intelligence(穆罕默德·本·扎耶德人工智能大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出任务向量组合方法,通过正交化分离特定价值偏好,实现语言模型伦理偏好对齐,并验证其跨语言有效性。
AI 中文摘要
大型语言模型(LLMs)日益部署于必须权衡相互冲突的道德价值观的应用中,然而即使是强大的模型也表现出隐藏的偏见,并在跨语言环境下指令遵循能力脆弱。我们引入了一个包含12,000个实例的两难选择数据集,涵盖三对价值观冲突:诚实与正义、正义与自主、自主与诚实,以及这些冲突的印地语、阿拉伯语、西班牙语和中文翻译,以探究跨语言行为。在GPT-5-mini上的基准测试显示,在没有给定策略的情况下,它在所有五种语言中始终倾向于选择诚实而非自主。Llama-3.2-1/3B模型表现出强烈的首选选项偏见;然而,普通微调和直接偏好优化(DPO)微调都能有效消除这种偏见,将准确率提高到98%以上。为了将数据集中学习相关性的效应与抽象价值观解耦,我们提出了一种基于任务向量迁移的实验,在计算某个价值偏好方向的任务向量后,我们将其与一般指令遵循向量正交化。我们的实验表明,该方法能有效隔离特定价值偏好的方向,并可成功用于任务算术,以获得具有相反立场的模型。
英文摘要
Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.
CommentsAccepted at the Pluralistic Alignment Workshop @ ICML 2026, Seoul, South Korea. https://icml.cc/virtual/2026/75692