AI 中文总结
针对联邦在线策略蒸馏中聚合-轨迹反馈导致的学习率敏感问题,提出FedTOPS方法,通过复用教师反馈调整更新幅度,在六个数学推理基准上使宏观Avg@8提升4.56至14.57个百分点。
AI 中文摘要
在线策略蒸馏(OPD)是一种有前景的语言模型适配方法,它使教师监督与学生自身生成的轨迹保持一致。当适配提示分布在多个客户端时,这一过程能否从联邦协作中获益?我们研究了联邦在线策略蒸馏,发现显著的合作收益可能被学习率敏感性所掩盖:在较小的学习率下,FedAvg的表现可能不优于独立的本地训练,而在较大的学习率下则能恢复明显的优势。我们通过学生作为学习者和未来训练数据生成者的双重角色来解释这一现象。聚合引起的优化滞后可能延迟获得有用的教师监督,进而减缓后续学习。我们的理论在一个具有共同最优解和稳定更新的可解模型中确立了这种聚合-轨迹反馈机制,并识别出学习率的两个耦合作用:从当前监督中学习和到达未来监督。在这一分析的指导下,我们提出了FedTOPS(联邦教师引导的在线策略缩放),该方法复用当前轨迹上的教师反馈,在客户端预测变化约束下调整FedAvg更新幅度。在六个数学推理基准上,FedTOPS在评估的学生模型和本地学习率下,相较于FedAvg将宏观Avg@8提升了4.56至14.57个百分点。
英文摘要
On-policy distillation (OPD) is a promising approach to language-model adaptation, aligning teacher supervision with the student's own generated trajectories. When adaptation prompts are distributed across clients, can this process benefit from federated collaboration? We study federated OPD and find that substantial collaboration gains can be obscured by learning-rate sensitivity: FedAvg can perform no better than independent local training at a small learning rate, yet recover a clear advantage at a larger rate. We explain this phenomenon through the student's dual role as learner and generator of future training data. An aggregation-induced optimization lag can delay access to useful teacher supervision, which in turn slows subsequent learning. Our theory establishes this aggregation--rollout feedback in a solvable model with a common optimum and stable updates, and identifies two coupled roles of learning rate: learning from current supervision and reaching future supervision. Guided by this analysis, we propose FedTOPS (Federated Teacher-guided On-Policy Scaling), which reuses teacher feedback on current trajectories to adapt the FedAvg update magnitude under clientwise predictive-change constraints. Across six mathematical reasoning benchmarks, FedTOPS improves macro Avg@8 over FedAvg by 4.56--14.57 percentage points across the evaluated student models and local learning rates.