arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于格的去中心化联邦学习中基于声誉的合作:演化博弈论方法

Reputation-driven Cooperation in Lattice-based Decentralized Federated Learning through Evolutionary Game Theory

Phuc Hoang Truong Huynh, Dung Tran Vinh, Khoa Duc Anh Lam, An Nghiem Nguyen Truong, Uyen Nha Tran Bui, Khang Nguyen Dinh, Bao Nguyen Le Gia, Minh Le Nguyen Nhat, Manh Hong Duong, The Anh Han, Thi Ai Thao Nguyen, and Le Hong Trang

arXiv 2608.01197首次发表:更新:

发表机构

Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT); Vietnam National University Ho Chi Minh City; School of Mathematics, University of Birmingham; School of Computing, Engineering and Digital Technologies, Teesside University(胡志明市技术大学计算机科学与工程学院; 越南国家大学胡志明市分校; 伯明翰大学数学学院; 提赛德大学计算、工程与数字技术学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对基于格的去中心化联邦学习的机会主义行为,本文提出结合有限理性、收益矩阵与声誉奖惩机制的演化博弈框架,仿真显示其在准确率、合作频率等指标上显著优于基线。

AI 中文摘要

去中心化联邦学习(Decentralized Federated Learning,DFL)已成为最优的隐私保护解决方案,但由于缺乏中央协调器,它仍易受机会主义行为影响。演化博弈论(Evolutionary Game Theory,EGT)是分析此类行为的强大框架,不过现有研究常假设智能体具有完全理性且策略固定。为解决这些局限,本文提出一种新颖的EGT框架,用于分析策略演化并提升整体系统性能。本文主要贡献有三:其一,在有限理性假设下,对格网络结构上的点对点(Peer-to-Peer,P2P)交互进行建模;其二,构建包含训练成本、通信开销与合作奖励的综合收益矩阵,同时定制策略更新规则以捕捉空间传播动态;其三,融入基于声誉的奖惩机制,有效遏制搭便车行为。仿真结果显示,该框架显著优于基线:平均准确率从约70%提升至82%,合作频率接近100%(基线低于5%),准确率方差从约0.40降至0.002,从而加快均匀收敛并保障系统稳定性。

英文摘要

Decentralized Federated Learning (DFL) has emerged as an optimal privacy-preserving solution; however, it remains vulnerable to opportunistic behaviors due to the absence of a central coordinator. While Evolutionary Game Theory (EGT) serves as a powerful framework for analyzing such behaviors, existing studies often assume that agents possess perfect rationality and maintain static strategies. To address these limitations, this paper proposes a novel EGT framework designed to analyze strategic evolution and enhance overall system performance. The primary contributions of this work are threefold: First, we model peer-to-peer (P2P) interactions on a lattice network structure under the assumption of bounded rationality. Second, we formulate a comprehensive payoff matrix incorporating training costs, communication overhead, and cooperative rewards, while tailoring a strategy update rule that captures spatial propagation dynamics. Third, we integrate a reputation-based reward-and-punishment mechanism to effectively deter free-riding behaviors. Simulation results demonstrate that the framework significantly outperforms the baseline. Specifically, it increases average accuracy from approximately 70% to 82%, elevates cooperation frequency to approach 100% (compared to below 5% in the baseline), and drops accuracy variance from around 0.40 to 0.002, thereby accelerating uniform convergence and ensuring system stability.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑