哪些目标需要拨盘?预测目标冲突与覆盖可操控多元对齐中的权衡
Which Objectives Need a Dial? Predicting Objective Conflict and Covering Trade-offs in Steerable Pluralistic Alignment
- University of Stuttgart(斯图加特大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对多元对齐中目标冲突与权衡覆盖问题,利用MODPO和预训练测量指标预测目标关系,并探索模型选择与参数合并策略,为构建可操控模型提供指导。
中文摘要 AI 辅助
人们持有多样且有时相互冲突的价值观,因此没有单一的、对齐的模型能够满足所有人。多元对齐因此要求模型具备可操控性,能够以不同方式平衡相互竞争的目标。多目标直接偏好优化(MODPO)通过使用目标权重来覆盖一系列权衡。我们研究两个问题:何时一个模型能同时改进两个目标,以及如何在不针对每个目标单独训练模型的情况下覆盖多种权衡?在来自HelpSteer和UltraFeedback的七对目标上,两个预训练测量指标能够预测人类标注数据中目标是对齐还是冲突,但对于AI标注数据则不然,因为响应长度和重复性混淆了奖励模型分数。对于更广泛的权衡覆盖,选择最近训练的模型和合并模型参数都有帮助,但两者都无法始终与直接训练相匹配。这些发现为构建服务于多样化偏好的可操控模型提供了实用指导。
英文摘要
People hold diverse, sometimes conflicting values, so no single aligned model can satisfy everyone. Pluralistic alignment therefore calls for steerable models that can balance competing objectives differently. Multi-Objective Direct Preference Optimization (MODPO) does this by using an objective weight to span a continuum of trade-offs. We study two questions: when can one model improve two objectives simultaneously, and how can many trade-offs be covered without training a separate model for each? Across seven objective pairs from HelpSteer and UltraFeedback, two pre-training measurements predict whether objectives align or conflict for human-annotated data, but not for AI-annotated data, where response length and repetition confound reward-model scores. For broader trade-off coverage, selecting the nearest trained model and merging model parameters both help, but neither consistently matches direct training. These findings yield practical guidance for building steerable models that serve diverse preferences.