发表机构
City University of Hong Kong; University of Washington; The Hong Kong University of Science and Technology(香港城市大学; 华盛顿大学; 香港科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对数据不足下用户偏好异质导致的梯度冲突,提出近似帕累托最优性(APO)方法,通过分组和受控上升学习共享初始化,实现少样本个性化对齐,实验验证其有效性。
AI 中文摘要
现实世界中的用户对大型语言模型(LLM)响应的多个目标往往表现出高度异质的偏好。轻量级对齐器(aligner)可以根据个人偏好定制这些响应,但稀缺的用户特定反馈使得个性化训练变得困难。跨用户学习共享初始化可以支持少样本(few-shot)适应。然而,异质偏好和竞争性目标导致跨用户和用户内部的梯度冲突,阻碍了有效的初始化学习。这引出了一个核心问题:我们如何协作地学习对齐器初始化,以支持对多样化用户偏好的少样本适应?为回答此问题,我们提出了近似帕累托最优性(Approximate Pareto Optimality, APO)方法。我们首先将更新兼容的用户分组,以便他们的信息可以以较少干扰的方式组合。在每个组内,我们将梯度下降与受控上升相结合,以协调竞争性目标并朝向帕累托前沿上特定于偏好的点移动。这产生了一个接近组内用户最优解的初始化。然后,我们使用少样本局部适应的更新迭代地对其进行细化,使其对个性化更有效。此外,我们为单局部步协作更新建立了条件次优性界,并刻画了初始化误差如何影响后续的随机适应。在Fed-ChatbotPA和UltraFeedback上的实验表明,仅使用20个本地示例,我们的方法相比现有方法取得了一致的改进。
英文摘要
Real-world users often exhibit highly heterogeneous preferences over multiple objectives for LLM responses. A lightweight aligner can tailor these responses to individual preferences, but scarce user-specific feedback makes personalized training difficult. Learning shared initializations across users can support few-shot adaptation. However, heterogeneous preferences and competing objectives cause gradient conflicts across users and within each user, hindering effective initialization learning. This raises a central question: \textbf{how can we collaboratively learn aligner initializations that support few-shot adaptation to diverse user preferences?} To answer this question, we propose \textbf{A}pproximate \textbf{P}areto \textbf{O}ptimality (APO). We first group users whose updates are compatible, so that their information can be combined with less interference. Within each group, we combine gradient descent with controlled ascent to coordinate competing objectives and move towards preference-specific points on the Pareto front. This produces an initialization that is close to the optima of the users in the group. We then iteratively refine it using updates from few-shot local adaptation, making it more effective for personalization. Furthermore, we establish conditional suboptimality bounds for a one-local-step collaborative update and characterize how initialization error affects subsequent stochastic adaptation. Experiments on Fed-ChatbotPA and UltraFeedback show consistent improvements over existing methods using only 20 local examples.