Fed-GRPO:奖励信号驱动的联邦组相对策略优化
Fed-GRPO: Reward-Signal-Driven Federated Group Relative Policy Optimization
- The University of Hong Kong(香港大学)
- Sun Yat-sen University(中山大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文针对GRPO需集中数据的隐私问题,提出Fed-GRPO联邦框架,含三种奖励驱动机制,在数学推理任务上性能优于FedAvg,通信压缩达621倍且精度损失轻微。
AI中文摘要:
大型语言模型(LLM)经强化学习(RL)微调后展现出强大的推理能力,尤其通过组相对策略优化(GRPO)实现。然而,现有GRPO方法假设训练数据可集中访问,受隐私或监管约束,实际中难以满足。为此,本文提出Fed-GRPO,一种联邦GRPO训练框架,通过无需共享原始数据的协作推理训练解决隐私约束,利用GRPO训练中自然产生的奖励统计作为零成本信号,指导聚合、本地训练与通信。Fed-GRPO包含三种奖励信号驱动机制:(i)信号加权聚合,按客户端奖励标准差加权,优先选择学习信号更强的客户端;(ii)全局奖励校准,基于本地-全局奖励差距重新加权每提示目标,引导客户端关注自身相对弱点;(iii)自适应稀疏通信,根据客户端更新的信息量分配带宽。在数学推理任务上的大量实验表明,Fed-GRPO在所有联邦方法中性能最优,明显优于FedAvg,接近集中式训练性能,同时通信量无损减少32倍,在严格带宽预算下支持最高621倍压缩,仅出现轻微精度下降。代码可在该https URL获取。
英文摘要:
Large Language Models (LLMs) have shown strong reasoning capabilities when fine-tuned with reinforcement learning (RL), particularly through Group Relative Policy Optimization (GRPO). However, existing GRPO methods assume centralized access to training data, which may not hold in practice due to privacy or regulatory constraints. To this end, we propose Fed-GRPO, a federated GRPO training framework that addresses these privacy constraints by enabling collaborative reasoning training without sharing raw data, which leverages the reward statistics naturally produced during GRPO training as zero-cost signals to guide aggregation, local training, and communication. Fed-GRPO contains three reward-signal-driven mechanisms: (i) \emph{signal-weighted aggregation} that weights clients by their reward standard deviation, prioritizing clients with stronger learning signals; (ii) \emph{global reward calibration} that re-weights per-prompt objectives based on the local-global reward gap, steering each client toward its relative weaknesses; and (iii) \emph{adaptive sparse communication} that allocates bandwidth based on the informativeness of each client's update. Extensive experiments on mathematical reasoning tasks demonstrate that Fed-GRPO achieves the best performance among all federated methods, clearly outperforms FedAvg and approaches centralized training performance, while losslessly reducing communication by $32\times$ and supporting up to $621\times$ compression under tight bandwidth budgets with only graceful accuracy degradation. Our code is available at https://github.com/HKU-HealthAI/Fed-GRPO.