A Systematic Evaluation of Preference Aggregation in Federated RLHF for Pluralistic Alignment of LLMs
对联邦RLHF中偏好聚合的系统评估:为LLM的多元化对齐
机构 * Department of Electrical Engineering and Computer Science University of California, Irvine(电气工程与计算机科学系 加州大学伊文斯分校)
专题命中 偏好对齐 :alignment(title,abstract);RLHF(title,abstract);分类 cs.CL、cs.AI
AI总结 本文提出了一种自适应偏好聚合方法,通过动态调整权重提升联邦RLHF中LLM的公平性与对齐性能。
Comments This paper is accepted at the NeurIPS 2025 Workshop on Evaluating the Evolving LLM Lifecycle