发表机构
School of Electrical Engineering and Computer Science; Digital Futures; KTH Royal Institute of Technology(电气工程与计算机科学学院; 数字未来机构; 瑞典皇家理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述针对协同学习从欧氏数据扩展至图结构数据的领域,梳理了基础原理、图分布分类与算法框架,明确了开放挑战与研究方向。
AI 中文摘要
传统机器学习方法(即在单一位置收集数据、训练模型并执行推理)面临可扩展性和隐私性等根本性限制,限制了其适用性。为应对这些挑战,近期研究探索了协同学习方法,包括联邦学习和去中心化学习,其中各智能体在本地执行训练和推理,仅进行有限协作。大多数协同学习研究聚焦于具有规则网格状结构的欧氏数据(如图像、文本),但这些方法无法捕捉许多实际应用中的关系模式,而图是这类模式的最佳代表。图学习依赖消息传递机制在相连节点间传播信息,在智能体需交换信息的协作环境中,这一机制在概念上十分契合。然而,在协作环境中对图结构数据进行学习的机遇与挑战仍未得到充分探索。本综述对从欧氏数据到图结构数据的协同学习展开全面研究,旨在整合这一新兴领域。我们首先回顾欧氏数据协同学习的基础原理,沿学习有效性、效率和隐私保护三个核心维度进行梳理;随后将讨论扩展至图结构数据,引入图分布场景的分类体系,刻画相关统计异质性,并构建标准化问题表述与算法框架;最后系统梳理了该领域的开放挑战与有前景的研究方向。
英文摘要
The conventional approach to machine learning, that is, collecting data, training models, and performing inference in a single location, faces fundamental limitations, including scalability and privacy, that restrict its applicability. To address these challenges, recent research has explored collaborative learning approaches, including federated learning and decentralized learning, where individual agents perform training and inference locally, with limited collaboration. Most collaborative learning research focuses on Euclidean data with regular, grid-like structure (e.g., images, text). However, these approaches fail to capture the relational patterns in many real-world applications, best represented by graphs. Learning on graphs relies on message-passing mechanisms to propagate information between connected nodes, making it conceptually well-suited for collaborative environments where agents must exchange information. Yet, the opportunities and challenges of learning on graph-structured data in collaborative settings remain largely underexplored. This survey provides a comprehensive investigation of collaborative learning from Euclidean to graph-structured data, aiming to consolidate this emerging field. We begin by reviewing its foundational principles for Euclidean data, organizing them along three core dimensions: learning effectiveness, efficiency, and privacy preservation. We then extend the discussion to graph-structured data, introducing a taxonomy of graph distribution scenarios, characterizing associated statistical heterogeneities, and developing standardized problem formulations and algorithmic frameworks. Finally, we systematically identify open challenges and promising research directions.
Comments96 pages. Published in Transactions on Machine Learning Research (TMLR), March 2026, with Survey Certification
Journal refTransactions on Machine Learning Research, 2026