发表机构
Institute of Engineering Research, Korea University; Kim Jaechul Graduate School of AI, KAIST; Industrial Management Engineering, Korea University(韩国大学工程研究所; 韩国科学技术院金在哲人工智能研究生院; 韩国大学产业管理工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对大语言模型个性化奖励建模的偏好异质性问题,提出FedGD方法,通过组去偏的客户端采样抵消组不平衡影响,实现无需预先知晓潜在组的有效个性化。
AI 中文摘要
大语言模型正通过奖励建模越来越多地对齐人类偏好,但用户偏好数据具有敏感性,通常无法集中存储。联邦学习将此类数据保留在本地,同时学习一个共享的初始奖励模型,之后通过本地微调为每个客户端进行个性化适配。由于用户常对同一组响应分配相反标签,现有联邦方法通过聚类相似客户端并为每组训练一个奖励模型来处理偏好异质性,假设每组需要各自的初始化。我们证明该假设并非必要:在偏好组平衡的情况下,单一的FedAvg模型尽管初始准确率接近随机,仅需几次本地优化步骤后,就能超越为每个真实组单独训练的奖励模型。我们将此现象归因于共享初始化的平坦性:跨所有客户端的平均操作学习到更丰富的共享表征,可区分不同响应,同时抵消冲突的偏好方向,使模型接近决策边界,能够快速适配。当组不平衡时,该效果会被破坏,因为抵消变得不对称,导致少数客户端距离决策边界过远而无法恢复。基于这一观察,我们提出FedGD(组去偏联邦学习),该方法在联邦训练过程中发现潜在偏好组,并通过组去偏的客户端采样学习单一奖励模型。通过抵消组不平衡的影响,FedGD学习到仍具有高度适应性的初始化,无需预先知晓潜在组即可实现有效的个性化。
英文摘要
Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. Federated learning keeps such data local while learning a shared initial reward model, which is later personalized for each client through local fine-tuning. Because users often assign opposite labels to the same pair of responses, existing federated methods address preference heterogeneity by clustering similar clients and training one reward model per group, assuming that each group requires its own initialization. We show that this assumption is unnecessary. Under balanced preference groups, a single FedAvg model, despite starting at nearly random accuracy, surpasses reward models trained separately for each ground-truth group after only a few local optimization steps. We attribute this phenomenon to the flatness of the shared initialization: averaging across all clients learns richer shared representations that distinguish responses while canceling conflicting preference directions, leaving the model near a decision boundary that can be rapidly adapted. Group imbalance breaks this effect as the cancellation becomes asymmetric and leaves minority clients too far from the boundary to recover. Motivated by this observation, we propose FedGD (Federated Learning with Group Debiasing), which discovers latent preference groups during federated training and learns a single reward model using group-debiased client sampling. By counteracting the effect of group imbalance, FedGD learns an initialization that remains highly adaptable, enabling effective personalization without prior knowledge of the underlying groups.