arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06896cs.RO

分布式安全学习控制:面向隐蔽执行器攻击下的大规模多机器人系统

Distributed Secure Learning Control for Large-scale Multirobots under Stealthy Actuator Attacks

Xinglong Zhang, Qingwen Ma, Cong Li, Hui Yin, Changxin Zhang, Yueying Wang, Wei Pan, Xin Xu

首次发表
浏览论文内容

中文总结 AI 辅助

针对大规模多机器人在隐蔽执行器攻击下的安全控制问题,提出分布式安全学习控制框架,结合强化学习与分布式模型预测控制,采用博弈论架构在线学习攻防策略,经仿真与实验验证有效。

中文摘要 AI 辅助

多机器人系统(MRS)的分布式学习控制在存在不确定性的情况下提供了显著的灵活性,但缺乏可证明的性能保证。一个有前景的方向是将强化学习(RL)集成到分布式模型预测控制(DMPC)中,利用RL在非线性策略设计中的优势以及DMPC的滚动时域重规划能力。然而,在恶意网络攻击(尤其是隐蔽攻击)下确保此类学习框架内的安全控制仍然是一个关键挑战,因为分布式策略的生成依赖于邻居之间的信息交换,而受损的智能体可以通过通信网络迅速影响其他智能体的行为。本文针对遭受恶意、隐蔽执行器攻击的大规模MRS,提出了一种分布式安全学习控制(DSLC)框架。我们的框架提供两个关键特性:(i)一种统一的方法,可在各种协调场景下实现安全学习控制;(ii)一种基于博弈论的分布式学习预测控制策略,通过基于微分博弈的DMPC框架学习如何平衡攻击者与防御者。具体而言,DSLC采用分布式攻击者-行动者-评论家架构,在每个预测区间内在线学习最优防御和攻击策略。与计算开环控制序列的基于数值优化的控制器不同,我们的方法以解析闭环形式同时生成对抗性攻击策略和相应的防御策略。防御策略可直接推广到具有不同规模和多样化执行器攻击概率的MRS。通过多种控制任务下的综合仿真和真实世界多轮式机器人实验,验证了DSLC的有效性和可扩展性。

英文摘要

Distributed learning control for multirobot systems (MRS) offers significant flexibility in presence of uncertainties but lacks provable performance guarantees. A promising direction involves integrating reinforcement learning (RL) into distributed model predictive control (DMPC), leveraging the strengths of RL in nonlinear policy design and the receding-horizon replanning capabilities of DMPC. However, ensuring secure control within such a learning framework under malicious cyber attacks, particularly stealthy ones, remains a critical challenge, because the distributed policies generation depends on information exchange among neighbors, where compromised agents can rapidly influence the behavior of others through the communication network. This article proposes a distributed secure learning control (DSLC) framework for large-scale MRS under malicious, stealthy actuator attacks. Our framework offers two key features: (i) a unified approach that enables secure learning control across various coordination scenarios and (ii) a game-theoretic distributed learning-based predictive control strategy that learns how to balance the attacker and defender through a differential-game based DMPC framework. Specifically, DSLC employs a distributed attacker-actor-critic architecture to learn the optimal defense and attack policies online within each prediction interval. Unlike numerical optimization-based controllers that calculate open-loop control sequences, our method simultaneously generates adversarial attack policies and corresponding defense policies in analytical closed-loop form. The defense policies could be directly generalized to MRS with varying scales and diverse actuator attack probabilities. The effectiveness and scalability of DSLC are validated through comprehensive simulations and real-world experiments in multiple wheeled robots via various control tasks.

发表机构

  • College of Intelligence Science and Technology, National University of Defense Technology(国防科技大学智能科学学院)
  • College of Mechanical and Vehicle Engineering, Hunan University(湖南大学机械与运载工程学院)
  • Department of Computer Science, The University of Manchester(曼彻斯特大学计算机科学系)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑