arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

异步P2P闲聊学习网络中知识蒸馏的收敛理论

Convergence Theory of Knowledge Distillation in Asynchronous P2P Gossip Learning Network

Lucas Qingyang Fang, Tiyao Liu, Jinhao Jing, Zeji Li, Kaijie Chen, Harikrishna Kuttivelil, Katia Obraczka

arXiv 2609.01952首次发表:更新:

发表机构

University of California, Santa Cruz; China University of Petroleum; The Chinese University of Hong Kong, Shenzhen; City University of Hong Kong; University of Virginia(加州大学圣克鲁兹分校; 中国石油大学; 香港中文大学(深圳); 香港城市大学; 弗吉尼亚大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对异步P2P闲聊学习网络,提出了完全去中心化KD的收敛理论,证明其可将函数分歧收缩40-61倍,为异构设备的无服务器学习提供了理论支撑。

AI 中文摘要

去中心化无服务器学习日益连接运行不同架构的设备,而标准工具去中心化SGD因无法对参数数量不同的模型求平均而无法使用。知识蒸馏(KD)交换软预测而非权重,避开了这一障碍,但完全去中心化异步P2P KD的收敛理论仍缺失。本文提供了该理论,将共识从参数空间转移到函数(输出)空间:KD事件是参考测度预测的希尔伯特空间中对等方预测分布上的logit空间几何收缩算子,在标准平滑性/方差假设以及两个可实现性假设(一个将参数SGD桥接到函数步,另一个控制受限任务/KD对齐)下,时间平均函数平稳性和函数空间分歧以O(1/(ηT))的速率收敛到O(η)+O(B_f²)+O(ζ_f²)邻域。其中B_f是任务最优解到对等方可达类的距离,ζ_f衡量持续的局部任务异质性。在同构、宽度异构和混合族网络的实验中,KD将函数分歧收缩了40-61倍,而孤立训练未做到;共享骨架主运行中,采样平稳性诊断的后期瞬态指数为0.99-1.90,四点步长扫描呈现出预测的瞬态邻域权衡。

英文摘要

Decentralized, serverless learning increasingly connects devices running different architectures, where the standard tool, decentralized SGD, is undefined as models with different parameter counts cannot be averaged. Knowledge distillation (KD) exchanges soft predictions rather than weights and sidesteps this obstacle, yet convergence theory for fully decentralized, asynchronous peer-to-peer (P2P) KD is lacking. We provide one, relocating consensus from parameter space to function (output) space: a KD event is a geometric contraction operator in logit space on the peers' predictive distributions, which we analyse in the Hilbert space of predictions on a reference measure. Under standard smoothness/variance assumptions and two realizability assumptions, one bridging parameter SGD to the functional step and one controlling restricted task/KD alignment, the time-averaged functional stationarity and function-space disagreement converge at rate $O(1/(ηT))$ to an $O(η)+O(B_f^2)+O(ζ_f^2)$ neighbourhood. Here $B_f$ is the distance from the task optimum to the peers' reachable classes and $ζ_f$ measures persistent local-task heterogeneity. Across homogeneous, width-heterogeneous, and mixed-family networks of the experiments, KD contracts function disagreement by $40-61\times$, while isolated training does not. The sampled stationarity diagnostic has late transient exponents $0.99-1.90$ on the shared-skeleton main runs, and the four-point step-size sweep exhibits the predicted transient: neighbourhood tradeoff.

Comments35 pages, 21 graphes, currently submitting to AAAI 2027 main track (Federated Learning and Decentralized Learning)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑