arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.11523cs.LG

何时进行干预?联邦强化学习中的状态感知稀疏操纵

When to Intervene? State-Aware Sparse Manipulation in Federated Reinforcement Learning

  • Sun Yat-sen University(中山大学)
  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

Shutong Zheng, Sijia Chen

AI总结:

本文针对联邦强化学习的拜占庭操纵问题,提出V-BSA攻击方法,通过状态感知稀疏干预实现高效攻击,揭示干预时机是序贯鲁棒性的独立维度。

AI中文摘要:

联邦强化学习(FRL)使分布式智能体能够协同训练决策策略,但其去中心化训练过程也使全局策略学习易受拜占庭操纵影响。现有投毒攻击主要聚焦于如何构建恶意更新,而轨迹级干预时机在很大程度上仍未明确。然而在序贯决策中,干预的应用位置会改变后续轨迹和学习信号。通过控制实验发现,即便恶意更新构建固定,改变所选轨迹状态也会显著改变攻击效能。因此,本文将“何时干预”确定为一个独立的攻击维度,并提出了可行性约束行为引导攻击(V-BSA),该方法利用局部策略不确定性选择稀疏干预状态,并应用包络约束行为引导。在离散动作基准测试中,V-BSA仅使用密集投毒所用干预的一小部分,就对鲁棒聚合器和集成防御实现了显著性能下降,同时揭示了依赖于任务和聚合方式的边界。总体而言,研究结果强调干预时机是FRL中序贯鲁棒性的一个独立维度,代码可在指定URL获取。

英文摘要:

Federated reinforcement learning (FRL) enables distributed agents to collaboratively train decision-making policies, but its decentralized training process also exposes global policy learning to Byzantine manipulation. Existing poisoning attacks primarily focus on how to construct malicious updates, while trajectory-level intervention timing remains largely implicit. In sequential decision making, however, where an intervention is applied can alter subsequent trajectories and learning signals. Through controlled experiments, we find that changing the selected trajectory states materially alters attack efficacy even when the malicious-update construction is fixed. We therefore identify when as a distinct attack dimension and introduce the Viability-constrained Behavioral Steering Attack (V-BSA), which uses local policy uncertainty to select sparse intervention states and applies envelope-constrained behavioral steering. Across discrete-action benchmarks, V-BSA achieves substantial degradation against robust aggregators and ensemble defenses with only a fraction of the interventions used by dense poisoning, while revealing task- and aggregation-dependent boundaries. Overall, our results highlight intervention timing as a distinct dimension of sequential robustness in FRL. The code is available at https://github.com/Yodeesy/V-BSA

补充信息

↑