训练哪种偏好?基于对抗偏好分布的端到端多目标对齐
Which Preferences to Train On? End-to-End Multi-Objective Alignment with an Adversarial Preference Distribution
- POSTECH(浦项科技大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对多目标对齐中偏好分布选择问题,提出MAESTRO方法,通过对抗性偏好分布的端到端RL训练单一模型,在多个任务上以最低成本实现最佳帕累托前沿。
AI中文摘要:
将大型语言模型(LLMs)与人类价值观对齐对于安全、高效且有益的AI部署至关重要。然而,人类价值观是多方面的:有用性、无害性和幽默感之间相互权衡,不同的用户需要不同的权衡。多目标对齐(MOA)通过训练一个能够提供帕累托前沿上任意点的策略来解决这一问题,但现有方法要么为每种偏好训练一个模型,要么事后对几个单独对齐的专家进行插值,要么训练一个单一的条件模型而不考虑应该训练哪些偏好。由于偏好单纯形的困难区域取决于当前的目标,现有方法使这些区域训练不足,未能充分利用单一模型。因此,我们提出了MAESTRO(通过端到端转向和鲁棒优化实现多目标对齐),它将MOA表述为偏好分布上的极小极大问题,并通过RL针对对抗性偏好分布端到端训练一个单一提示条件策略:该分布是通过在线镜像下降更新的狄利克雷分布,朝向当前策略服务最差的偏好,而不是固定的分布。在HH-RLHF、BeaverTails和摘要任务上,最多包含三个目标,MAESTRO在单次训练运行中在大多数任务上取得了最佳帕累托前沿,且训练成本在比较方法中最低。最大的差距出现在固定偏好分布训练不足的困难区域,这证实了单一提示条件模型能够自行覆盖目标权衡。
英文摘要:
Aligning large language models (LLMs) with human values is important for safe, efficient, and beneficial AI deployment. However, human values are multifaceted: helpfulness, harmlessness and humor trade off against one another, and different users want different trade-offs. Multi-objective alignment (MOA) addresses this by training a policy that can provide any point of the Pareto front, but existing methods either train one model per preference, interpolate a few separately aligned experts post hoc, or train a single conditioned model without considering which preferences it should be trained on. Since the hard regions of the preference simplex depend on the objectives at hand, existing methods leave them under-trained and do not get the most out of a single model. Therefore, we propose MAESTRO (Multi-objective Alignment via End-to-end STeering and Robust Optimization), which formulates MOA as a minimax problem over preference distributions and trains a single prompt-conditioned policy end-to-end with RL against an adversarial preference distribution: a Dirichlet distribution updated by online mirror descent toward the preferences the current policy serves worst, rather than on a fixed one. On HH-RLHF, BeaverTails and a summarization task, with up to three objectives, MAESTRO attains the best Pareto front on most tasks in a single training run, at the lowest training cost among the compared methods. The largest margins appear in the hard regions that a fixed preference distribution leaves under-trained, confirming that a single prompt-conditioned model is capable of covering the objective trade-offs on its own.